使用统计学理解数据
使用 Python 中的统计学理解数据
Section titled “使用 Python 中的统计学理解数据”在机器学习中,数据是基础。模型的质量很大程度上取决于数据的质量和理解程度。在将数据输入任何算法之前,了解其特征至关重要。
数据可视化提供图形洞察(在另一章节介绍),而描述性统计(Descriptive statistics)提供数据属性的定量摘要。这种初步分析有助于识别潜在问题,如异常值(Outliers)、缺失值(Missing values)或偏态分布(Skewed distributions),从而指导数据预处理(Preprocessing)步骤。
本章重点介绍如何使用 Python,特别是 Pandas 库,对数据集执行基本的统计分析。
预览原始数据
Section titled “预览原始数据”第一步总是查看实际数据值。这能让你对数据集有一个定性感官认识,并能立即发现诸如格式不一致或意外值等问题。
Pandas 中的 head() 和 tail() 方法非常适合此目的。我们加载 Pima Indian Diabetes 数据集(一个常用于二分类 Binary classification 的数据集)并查看前几行。
注意:确保已安装 ‘pandas’ 和 ‘scikit-learn’ 库(pip install pandas scikit-learn)。我们将使用 scikit-learn 的实用工具轻松获取数据集,避免硬编码文件路径。
import pandas as pdfrom sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes dataset from OpenML# Note: fetch_openml might require an internet connection the first timepima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')df = pima.frame
# Define more descriptive column names (based on dataset description)df.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
# Display the first 10 rowsprint("First 10 rows of the dataset:")print(df.head(10))
# Display the last 5 rowsprint("\nLast 5 rows of the dataset:")print(df.tail(5))输出(示例)
Section titled “输出(示例)”First 10 rows of the dataset: num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age class0 6.0 148.0 72.0 35.0 0.0 33.6 0.627 50 11 1.0 85.0 66.0 29.0 0.0 26.6 0.351 31 02 8.0 183.0 64.0 0.0 0.0 23.3 0.672 32 13 1.0 89.0 66.0 23.0 94.0 28.1 0.167 21 04 0.0 137.0 40.0 35.0 168.0 43.1 2.288 33 15 5.0 116.0 74.0 0.0 0.0 25.6 0.201 30 06 3.0 78.0 50.0 32.0 88.0 31.0 0.248 26 17 10.0 115.0 0.0 0.0 0.0 35.3 0.134 29 08 2.0 197.0 70.0 45.0 543.0 30.5 0.158 53 19 8.0 125.0 96.0 0.0 0.0 0.0 0.232 54 1
Last 5 rows of the dataset: num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age class763 10.0 101.0 76.0 48.0 180.0 32.9 0.171 63 0764 2.0 122.0 70.0 27.0 0.0 36.8 0.340 27 0765 5.0 121.0 72.0 23.0 112.0 26.2 0.245 30 0766 1.0 126.0 60.0 0.0 0.0 30.1 0.349 47 1767 1.0 93.0 70.0 31.0 0.0 30.4 0.315 23 0观察前几行可以发现数据类型(多数为数值型)并大致了解数值范围。请注意,像 ‘blood_pressure’、‘skin_thickness’、‘insulin’ 和 ‘bmi’ 等一些列存在零值,这可能表示这些生理测量值的缺失数据,需要进一步调查或填充(Imputation)。
检查数据维度
Section titled “检查数据维度”了解数据集的大小(行数/样本数和列数/特征数)对于规划至关重要:
- 特征过多可能导致“维度灾难”(Curse of dimensionality),需要进行特征选择或特征工程。
- 样本过少可能限制模型的泛化能力。
- 大型数据集需要更多计算资源,可能需要不同的算法方法(例如,增量学习 Incremental learning、分布式计算 Distributed computing)。
Pandas DataFrame 的 shape 属性提供了这些信息。
import pandas as pdfrom sklearn.datasets import load_iris
# Load the Iris datasetiris = load_iris(as_frame=True)df_iris = iris.frame
# Add the target column for completeness (optional for shape)df_iris['target'] = iris.target
print(f"Dataset dimensions (rows, columns): {df_iris.shape}")Dataset dimensions (rows, columns): (150, 5)输出表明 Iris 数据集有 150 个样本和 5 列(4 个特征 + 1 个目标变量)。
理解属性数据类型
Section titled “理解属性数据类型”机器学习算法通常需要特定的数据类型。例如,大多数算法期望数值输入。理解数据类型(dtypes)有助于识别可能需要转换的列(例如,将分类字符串转换为数值表示,如 One-Hot 编码 One-hot encoding 或 Label 编码 Label encoding)。
Pandas 在加载时会自动推断数据类型,但使用 dtypes 属性或 info() 方法(它也显示非空计数)进行验证是良好的实践。
import pandas as pdfrom sklearn.datasets import load_iris
# Load the Iris datasetiris = load_iris(as_frame=True)df_iris = iris.framedf_iris['target'] = iris.target
print("Data types of each column:")print(df_iris.dtypes)
print("\n--- More Detailed Info ---")df_iris.info()Data types of each column:sepal length (cm) float64sepal width (cm) float64petal length (cm) float64petal width (cm) float64target int64dtype: object
--- More Detailed Info ---<class 'pandas.core.frame.DataFrame'>RangeIndex: 150 entries, 0 to 149Data columns (total 5 columns): # Column Non-Null Count Dtype--- ------ -------------- ----- 0 sepal length (cm) 150 non-null float64 1 sepal width (cm) 150 non-null float64 2 petal length (cm) 150 non-null float64 3 petal width (cm) 150 non-null float64 4 target 150 non-null int64dtypes: float64(4), int64(1)memory usage: 6.0 KB输出确认所有特征都是 float64,目标变量是 int64,适用于大多数数值算法。info() 输出也确认在这个特定的数据集中没有缺失值。
描述性统计摘要
Section titled “描述性统计摘要”除了维度和类型,对每个数值属性的中心趋势(Central tendency)、离散度(Dispersion)和分布形状的摘要也至关重要。Pandas 中的 describe() 方法提供了关键统计量:
- Count:非空观测值数量。
- Mean:均值(平均值)。
- Std:标准差(Standard Deviation),衡量围绕均值的离散度。
- Min:最小值。
- 25%:第一四分位数(Q1)。
- 50%:中位数(Median,Q2)。
- 75%:第三四分位数(Q3)。
- Max:最大值。
这个摘要有助于识别不同属性的尺度(可能需要进行特征缩放 Scaling/Normalization)、异常值(中位数和均值之间的巨大差异,或极端的最小值/最大值)的存在,以及基本分布。
import pandas as pdfrom sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes dataset from OpenMLpima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')df = pima.framedf.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
# Set display options for better readabilitypd.set_option('display.width', 100)pd.set_option('display.precision', 3)
print("Dataset dimensions:")print(df.shape)
print("\nStatistical Summary:")print(df.describe())Dataset dimensions:(768, 9)
Statistical Summary: num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age classcount 768.000 768.000 768.000 768.000 768.000 768.000 768.000 768.000 768.000mean 3.845 120.895 69.105 20.536 79.799 31.993 0.472 33.241 0.349std 3.370 31.973 19.356 15.952 115.244 7.884 0.331 11.760 0.477min 0.000 0.000 0.000 0.000 0.000 0.000 0.078 21.000 0.00025% 1.000 99.000 62.000 0.000 0.000 27.300 0.244 24.000 0.00050% 3.000 117.000 72.000 23.000 30.500 32.000 0.372 29.000 0.00075% 6.000 140.250 80.000 32.000 127.250 36.600 0.626 41.000 1.000max 17.000 199.000 122.000 99.000 846.000 67.100 2.420 81.000 1.000摘要确认有 768 个样本。再次观察到 ‘glucose_concentration’、‘blood_pressure’、‘skin_thickness’、‘insulin’ 和 ‘bmi’ 等列的最小值为 0。这再次证实了将零值编码为缺失值的猜测,需要谨慎处理(例如,数据填充 Imputation)。
查看类别分布(用于分类)
Section titled “查看类别分布(用于分类)”在分类任务中,理解目标变量(类别标签)的分布至关重要。不平衡类别(Imbalanced classes,指数据集中不同类别的样本数量差异显著的情况)可能使模型偏向于多数类别,可能需要特定的技术,如重采样(Resampling,过采样 Oversampling/欠采样 Undersampling)或使用不同的评估指标(例如,F1 分数 F1-score、AUC)。
我们可以使用 Pandas 的 groupby() 方法,然后结合 size() 来计算每个类别的出现次数。
import pandas as pdfrom sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes datasetpima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')df = pima.framedf.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
# Convert class column to integer if it's not already (fetch_openml might load it as categorical)df['class'] = pd.to_numeric(df['class'])
# Calculate class distributionclass_counts = df.groupby('class').size()
print("Class Distribution:")print(class_counts)Class Distribution:class0 5001 268dtype: int64输出显示类别 0(无糖尿病)有 500 个样本,类别 1(有糖尿病)有 268 个样本。虽然不算极端不平衡,但差异很明显,在建模时可能需要考虑。
查看属性之间的相关性
Section titled “查看属性之间的相关性”相关性(Correlation)衡量数值属性对之间的线性关系。Pearson 相关系数(Pearson correlation coefficient)是常用的指标:
- 值接近 +1:强正线性相关。
- 值接近 -1:强负线性相关。
- 值接近 0:弱或无线性相关。
识别高度相关的特征很重要,因为:
- 冗余(Redundancy): 高度相关的特征可能提供冗余信息。
- 模型稳定性(Model Stability): 某些算法(如线性回归 Linear Regression)在面对高度相关的预测变量(多重共线性 Multicollinearity)时可能变得不稳定。 可能需要进行特征选择或降维技术(如 PCA)。
Pandas 的 corr() 方法计算成对相关矩阵。
import pandas as pdfrom sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes datasetpima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')df = pima.framedf.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']df['class'] = pd.to_numeric(df['class'])
# Set display optionspd.set_option('display.width', 100)pd.set_option('display.precision', 3)
# Calculate correlation matrixcorrelations = df.corr(method='pearson')
print("Correlation Matrix:")print(correlations)Correlation Matrix: num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age classnum_pregnancies 1.000 0.129 0.141 -0.082 -0.074 0.018 -0.034 0.544 0.222glucose_concentration 1.000 1.000 0.153 0.057 0.331 0.221 0.137 0.264 0.467blood_pressure 0.141 0.153 1.000 0.207 0.089 0.282 0.041 0.240 0.065skin_thickness -0.082 0.057 0.207 1.000 0.437 0.393 0.184 -0.114 0.075insulin -0.074 0.331 0.089 0.437 1.000 0.198 0.185 -0.042 0.131bmi 0.018 0.221 0.282 0.393 0.198 1.000 0.141 0.036 0.293diabetes_pedigree -0.034 0.137 0.041 0.184 0.185 0.141 1.000 0.034 0.174age 0.544 0.264 0.240 -0.114 -0.042 0.036 0.034 1.000 0.238class 0.222 0.467 0.065 0.075 0.131 0.293 0.174 0.238 1.000该矩阵显示了所有属性对之间的相关性。例如,‘age’ 和 ‘num_pregnancies’ 具有中等程度的正相关性(0.544)。‘glucose_concentration’ 与目标变量 ‘class’ 的相关性最高(0.467),在所示特征中。
查看属性分布的偏度
Section titled “查看属性分布的偏度”偏度(Skewness)衡量概率分布的不对称性。许多机器学习算法(特别是线性模型和假设正态性的算法)在特征呈对称(类似高斯分布)分布时表现更好。
- 正偏态(Positive Skew,右偏):尾部在右侧更长。均值 > 中位数。
- 负偏态(Negative Skew,左偏):尾部在左侧更长。均值 < 中位数。
- 零偏态(Zero Skew):对称分布(如完美的正态分布 Gaussian)。
显著的偏度可能需要在预处理过程中进行数据转换(例如,对数 Log、平方根 Square root、Box-Cox 转换)。Pandas 的 skew() 方法计算每个数值列的偏度。
import pandas as pdfrom sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes datasetpima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')df = pima.framedf.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']df['class'] = pd.to_numeric(df['class'])
# Calculate skewnessskewness = df.skew()
print("Skewness of each attribute:")print(skewness)Skewness of each attribute:num_pregnancies 0.902glucose_concentration 0.174blood_pressure -1.844skin_thickness 0.109insulin 2.272bmi -0.429diabetes_pedigree 1.920age 1.130class 0.635dtype: float64接近 0 的值表示低偏度(例如,‘glucose_concentration’、‘skin_thickness’)。与 0 显著不同的值表示偏度。例如,‘insulin’ (2.272) 和 ‘diabetes_pedigree’ (1.920) 显示显著的正偏态,而 ‘blood_pressure’ (-1.844) 显示显著的负偏态(可能受到零值的影响)。这些属性可能从转换中受益。
- Pandas 文档:https://pandas.pydata.org/pandas-docs/stable/
- Scikit-learn 数据集:https://scikit-learn.org/stable/datasets.html
- OpenML 数据集:https://www.openml.org/