Skip to content

Scikit Learn - 建模过程

本章涵盖了 Scikit-learn 中典型的模型构建流程。我们将探索每个步骤,从数据集加载和准备开始,然后转向模型训练、评估和持久化。理解这个工作流程对于有效应用机器学习技术至关重要。

数据集是数据的集合,通常组织成特征 (features) 和目标 (target)(或响应,response)。关键组成部分包括:

Features (X):这些是模型用于进行预测的输入变量或属性。也称为预测变量 (predictors) 或自变量 (independent variables)。

  • 特征矩阵 (Feature Matrix):一种二维数组状结构(例如,NumPy 数组或 Pandas DataFrame),其中行代表样本 (samples),列代表特征 (features)。形状为:[n_samples, n_features]。
  • 特征名称 (Feature Names):对应于特征矩阵中每列名称的列表,有助于理解数据。

Response (y):这是模型旨在预测的输出变量。也称为目标 (target)、标签 (label) 或因变量 (dependent variable)。

  • 响应向量 (Response Vector):一种一维数组状结构,代表每个样本的目标值。形状为:[n_samples]。
  • 目标名称 (Target Names):对于分类任务,这代表了响应向量可以取的不同类别 (classes) 的名称。

Scikit-learn 包含几个内置数据集供练习使用,例如用于分类 (classification) 的 Iris 和 Digits 数据集,以及用于回归 (regression) 的 California Housing 数据集。这些数据集对于快速测试算法非常有用。您可以在 sklearn.datasets 模块下找到更多。

以下是加载经典的 Iris 数据集的方法。我们使用 return_X_y=True 以便直接获取特征 (X) 和目标 (y)。

from sklearn.datasets import load_iris
# Load the dataset
iris = load_iris()
X = iris.data
y = iris.target
# Alternatively, load features and target directly
# X, y = load_iris(return_X_y=True)
feature_names = iris.feature_names
target_names = iris.target_names
print(f"Feature names: {feature_names}")
print(f"Target names: {target_names}")
print(f"\nFirst 5 rows of X (Features):\n{X[:5]}")
print(f"\nFirst 5 labels of y (Target):\n{y[:5]}")
print(f"\nShape of X: {X.shape}")
print(f"Shape of y: {y.shape}")
Feature names: ['sepal length (cm)', 'sepal width (cm)', 'petal length (cm)', 'petal width (cm)']
Target names: ['setosa' 'versicolor' 'virginica']
First 5 rows of X (Features):
[[5.1 3.5 1.4 0.2]
[4.9 3. 1.4 0.2]
[4.7 3.2 1.3 0.2]
[4.6 3.1 1.5 0.2]
[5. 3.6 1.4 0.2]]
First 5 labels of y (Target):
[0 0 0 0 0]
Shape of X: (150, 4)
Shape of y: (150,)

为了评估模型在未见过数据上的性能,将数据集划分为训练集 (training set) 和测试集 (testing set) 至关重要。模型从训练集中学习,然后在测试集上评估其性能。这有助于防止过拟合 (overfitting),即模型在训练数据上表现良好,但在新数据上表现较差。

sklearn.model_selection 中的 train_test_split 函数是常用的方法。这里,我们划分 Iris 数据集,分配 30% 用于测试,70% 用于训练。

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
# Load data
X, y = load_iris(return_X_y=True)
# Split data: 70% training, 30% testing
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42, stratify=y
)
print(f"Shape of X_train: {X_train.shape}")
print(f"Shape of X_test: {X_test.shape}")
print(f"Shape of y_train: {y_train.shape}")
print(f"Shape of y_test: {y_test.shape}")
Shape of X_train: (105, 4)
Shape of X_test: (45, 4)
Shape of y_train: (105,)
Shape of y_test: (45,)

train_test_split 的关键参数:

  • X, y: 要划分的特征矩阵和响应向量。
  • test_size: 测试集的比例大小(例如,0.3 表示 30%)。也可以使用 train_size。
  • random_state: 用于随机数生成器的种子。设置此参数可确保每次运行代码时划分结果相同,这对于可重现性至关重要。可以使用任何整数值。
  • stratify=y:(建议用于分类任务)确保训练集和测试集中目标类别比例与原始数据集相似。这对于不平衡数据集很重要。

准备好数据后,下一步是训练一个机器学习模型。Scikit-learn 提供了多种算法,它们都遵循一致的 API:fit() 用于训练,predict() 用于进行预测,以及 score() 用于评估。

我们将使用 K 近邻 (KNN) 算法,这是一个简单但有效的分类器。此示例演示了基本的训练和预测工作流程。

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
import numpy as np
# 加载并划分数据(如前所示)
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42, stratify=y)
# 初始化 KNN 分类器
# n_neighbors 指定了 KNN 中的 'K'
knn_classifier = KNeighborsClassifier(n_neighbors=3)
# 使用训练数据训练模型
knn_classifier.fit(X_train, y_train)
# 在测试数据上进行预测
y_pred = knn_classifier.predict(X_test)
# 评估模型的准确率
accuracy = accuracy_score(y_test, y_pred)
print(f"Accuracy on test set: {accuracy:.4f}")
# 对新的、未见过的数据样本进行预测示例
# 注意:新的样本应具有与训练数据相同的特征数量(本例中为 4)
sample_new_data = np.array([
[5.0, 3.1, 1.6, 0.2], # 期望结果:setosa (0)
[6.5, 3.0, 5.2, 2.0] # 期望结果:virginica (2)
])
predicted_species_indices = knn_classifier.predict(sample_new_data)
iris = load_iris() # 重新加载以便获取 target_names
predicted_species_names = iris.target_names[predicted_species_indices]
print(f"Predictions for new samples: {predicted_species_names}")
Accuracy on test set: 0.9778
Predictions for new samples: ['setosa' 'virginica']

关于 KNN 的更多细节,请参阅其专门的章节。选择正确的模型和调整其参数(例如 KNN 的 n_neighbors)通常涉及交叉验证 (cross-validation) 等技术,这将在后面讨论。

训练模型可能很耗时。模型训练完成后,您通常希望保存(持久化,persist)它以便将来使用,而无需重新训练。joblib 库(SciPy 生态系统的一部分,并由 Scikit-learn 推荐)非常适合此目的,特别是对于包含大型 NumPy 数组的对象。

让我们保存之前训练好的 knn_classifier 模型:

import joblib
# 假设 knn_classifier 是我们上一步训练好的模型
# 将模型保存到文件
model_filename = 'iris_knn_classifier.joblib'
joblib.dump(knn_classifier, model_filename)
print(f"Model saved to {model_filename}")
# ... 稍后,在另一个脚本或会话中 ...
# 从文件加载模型
loaded_knn_model = joblib.load(model_filename)
print(f"Model loaded from {model_filename}")
# 现在可以使用加载的模型 loaded_knn_model 进行预测
# 示例:y_pred_loaded = loaded_knn_model.predict(X_test)
# print(f"加载模型的准确率: {accuracy_score(y_test, y_pred_loaded):.4f}")

这将保存整个模型对象,包括其学习到的参数。请确保加载模型的环境中具有相同版本的相关库(如 Scikit-learn),以保证兼容性。

原始数据通常不适合直接输入到机器学习算法中。预处理 (Preprocessing) 将原始数据转换为更可用的格式。Scikit-learn 的 sklearn.preprocessing 模块提供了各种工具。常见的预处理步骤包括缩放 (scaling)、归一化 (normalization) 和分类特征编码 (encoding categorical features)。

为什么需要预处理?许多算法(例如,SVM、KNN、PCA、神经网络)对特征尺度 (feature scales) 很敏感。具有较大值的特征可能会主导学习过程。预处理有助于标准化数据。

二值化 (Binarization) 根据阈值将数值特征转换为布尔值(0 或 1)。这对于从连续特征创建二元特征很有用。

import numpy as np
from sklearn.preprocessing import Binarizer
input_data = np.array([
[2.1, -1.9, 5.5],
[-1.5, 2.4, 3.5],
[0.5, -7.9, 5.6],
[5.9, 2.3, -5.8]
])
# 大于 0.5 的值变为 1,否则变为 0
binarizer = Binarizer(threshold=0.5)
data_binarized = binarizer.fit_transform(input_data)
print(f"\nBinarized data (threshold=0.5):\n{data_binarized}")
Binarized data (threshold=0.5):
[[1. 0. 1.]
[0. 1. 1.]
[0. 0. 1.]
[1. 1. 0.]]

去除均值,更常见地称为标准化 (standardization) 或 Z-score 归一化 (Z-score normalization),将特征转换为具有零均值和单位方差。这是通过减去均值并除以每个特征的标准差来实现的。

import numpy as np
from sklearn.preprocessing import StandardScaler
input_data = np.array([
[2.1, -1.9, 5.5],
[-1.5, 2.4, 3.5],
[0.5, -7.9, 5.6],
[5.9, 2.3, -5.8]
], dtype=np.float64)
print(f"Original Mean: {input_data.mean(axis=0)}")
print(f"Original Std Dev: {input_data.std(axis=0)}")
scaler = StandardScaler()
data_scaled = scaler.fit_transform(input_data)
print(f"\nScaled Data (StandardScaler):\n{data_scaled}")
print(f"Mean of scaled data: {data_scaled.mean(axis=0)}")
print(f"Std Dev of scaled data: {data_scaled.std(axis=0)}")
Original Mean: [ 1.75 -1.275 2.2 ]
Original Std Dev: [2.71431391 4.20022321 4.69414529]
Scaled Data (StandardScaler):
[[ 0.12896806 0.6318175 0.70299018]
[-1.19734452 0.87470814 0.2769086 ]
[-0.46049308 -1.57685687 0.72410327]
[ 1.52886954 0.07033123 -1.7039976 ]]
Mean of scaled data: [2.22044605e-17 0.00000000e+00 0.00000000e+00]
Std Dev of scaled data: [1. 1. 1.]

注意:关键在于只在训练数据上 fit 缩放器 (scaler),然后在训练数据和测试数据上都使用 transform,以防止测试集中的数据泄露 (data leakage)。

最小-最大缩放 (Min-Max scaling)(或归一化,normalization)将特征重新缩放到特定范围,通常是 [0, 1] 或 [-1, 1]。这是通过减去最小值并除以每个特征的范围(最大值 - 最小值)来完成的。

import numpy as np
from sklearn.preprocessing import MinMaxScaler
input_data = np.array([
[2.1, -1.9, 5.5],
[-1.5, 2.4, 3.5],
[0.5, -7.9, 5.6],
[5.9, 2.3, -5.8]
], dtype=np.float64)
# 将特征缩放到 [0, 1] 范围
min_max_scaler = MinMaxScaler(feature_range=(0, 1))
data_scaled_minmax = min_max_scaler.fit_transform(input_data)
print(f"\nMin-Max Scaled Data (range [0, 1]):\n{data_scaled_minmax}")
Min-Max Scaled Data (range [0, 1]):
[[0.48648649 0.58252427 0.99122807]
[0. 1. 0.81578947]
[0.27027027 0. 1. ]
[1. 0.99029126 0. ]]

Scikit-learn 中的归一化 (Normalization)(使用 sklearn.preprocessing.Normalizer)独立地重新缩放每个样本(行),使其具有单位范数 (unit norm)(L1 或 L2)。这与 StandardScaler 或 MinMaxScaler 等按特征(列)进行缩放不同。归一化通常用于文本分类或聚类。

重新缩放每个样本,使其特征绝对值之和为 1。

import numpy as np
from sklearn.preprocessing import Normalizer
input_data = np.array([
[2.1, -1.9, 5.5],
[-1.5, 2.4, 3.5],
[0.5, -7.9, 5.6],
[5.9, 2.3, -5.8]
], dtype=np.float64)
normalizer_l1 = Normalizer(norm='l1')
data_normalized_l1 = normalizer_l1.fit_transform(input_data)
print(f"\nL1 Normalized Data:\n{data_normalized_l1}")
print(f"Sum of absolute values per row (L1 norm): {np.sum(np.abs(data_normalized_l1), axis=1)}")
L1 Normalized Data:
[[ 0.22105263 -0.2 0.57894737]
[-0.2027027 0.32432432 0.47297297]
[ 0.03571429 -0.56428571 0.4 ]
[ 0.42142857 0.16428571 -0.41428571]]
Sum of absolute values per row (L1 norm): [1. 1. 1. 1.]

重新缩放每个样本,使其特征平方和(欧几里得范数,Euclidean norm)为 1。

import numpy as np
from sklearn.preprocessing import Normalizer
input_data = np.array([
[2.1, -1.9, 5.5],
[-1.5, 2.4, 3.5],
[0.5, -7.9, 5.6],
[5.9, 2.3, -5.8]
], dtype=np.float64)
normalizer_l2 = Normalizer(norm='l2')
data_normalized_l2 = normalizer_l2.fit_transform(input_data)
print(f"\nL2 Normalized Data:\n{data_normalized_l2}")
print(f"Sum of squares per row (L2 norm squared): {np.sum(data_normalized_l2**2, axis=1)}")
L2 Normalized Data:
[[ 0.33946114 -0.30713151 0.88906489]
[-0.33325106 0.53320169 0.7775858 ]
[ 0.05156558 -0.81473612 0.57753446]
[ 0.68706914 0.26784051 -0.6754239 ]]
Sum of squares per row (L2 norm squared): [1. 1. 1. 1.]

Scikit-learn 的 Pipeline 对象允许您将多个处理步骤(如缩放和 PCA)与最终的估计器 (estimator)(如分类器,classifier)串联起来。强烈建议使用它来实现:

  • 便捷性:将多个步骤组合成一个对象。
  • 防止数据泄露:确保在交叉验证期间仅在训练数据上拟合 (fit) 变换。
  • 可重现性:封装整个预处理和模型构建工作流程。

简单 Pipeline 示例:

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
# 加载并划分数据
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# 创建 Pipeline:StandardScaler 然后 SVC
pipe = Pipeline([
('scaler', StandardScaler()), # 步骤 1:缩放数据
('svc', SVC()) # 步骤 2:支持向量分类器
])
# 拟合 Pipeline(应用 scaler 的 fit_transform,然后应用 SVC 的 fit)
pipe.fit(X_train, y_train)
# 进行预测(应用 scaler 的 transform,然后应用 SVC 的 predict)
y_pred_pipeline = pipe.predict(X_test)
accuracy_pipeline = accuracy_score(y_test, y_pred_pipeline)
print(f"Accuracy with Pipeline: {accuracy_pipeline:.4f}")
Accuracy with Pipeline: 1.0000

Pipeline 是构建健壮且可维护的机器学习模型的强大工具。您可以在 Scikit-learn 关于 sklearn.pipeline.Pipeline 的文档中了解更多。