Scikit Learn - 建模过程
Scikit-learn - 模型构建流程
Section titled “Scikit-learn - 模型构建流程”本章涵盖了 Scikit-learn 中典型的模型构建流程。我们将探索每个步骤,从数据集加载和准备开始,然后转向模型训练、评估和持久化。理解这个工作流程对于有效应用机器学习技术至关重要。
数据集是数据的集合,通常组织成特征 (features) 和目标 (target)(或响应,response)。关键组成部分包括:
Features (X):这些是模型用于进行预测的输入变量或属性。也称为预测变量 (predictors) 或自变量 (independent variables)。
- 特征矩阵 (Feature Matrix):一种二维数组状结构(例如,NumPy 数组或 Pandas DataFrame),其中行代表样本 (samples),列代表特征 (features)。形状为:
[n_samples, n_features]。 - 特征名称 (Feature Names):对应于特征矩阵中每列名称的列表,有助于理解数据。
Response (y):这是模型旨在预测的输出变量。也称为目标 (target)、标签 (label) 或因变量 (dependent variable)。
- 响应向量 (Response Vector):一种一维数组状结构,代表每个样本的目标值。形状为:
[n_samples]。 - 目标名称 (Target Names):对于分类任务,这代表了响应向量可以取的不同类别 (classes) 的名称。
Scikit-learn 包含几个内置数据集供练习使用,例如用于分类 (classification) 的 Iris 和 Digits 数据集,以及用于回归 (regression) 的 California Housing 数据集。这些数据集对于快速测试算法非常有用。您可以在 sklearn.datasets 模块下找到更多。
示例:加载 Iris 数据集
Section titled “示例:加载 Iris 数据集”以下是加载经典的 Iris 数据集的方法。我们使用 return_X_y=True 以便直接获取特征 (X) 和目标 (y)。
from sklearn.datasets import load_iris
# Load the datasetiris = load_iris()X = iris.datay = iris.target
# Alternatively, load features and target directly# X, y = load_iris(return_X_y=True)
feature_names = iris.feature_namestarget_names = iris.target_names
print(f"Feature names: {feature_names}")print(f"Target names: {target_names}")print(f"\nFirst 5 rows of X (Features):\n{X[:5]}")print(f"\nFirst 5 labels of y (Target):\n{y[:5]}")print(f"\nShape of X: {X.shape}")print(f"Shape of y: {y.shape}")Feature names: ['sepal length (cm)', 'sepal width (cm)', 'petal length (cm)', 'petal width (cm)']Target names: ['setosa' 'versicolor' 'virginica']
First 5 rows of X (Features):[[5.1 3.5 1.4 0.2] [4.9 3. 1.4 0.2] [4.7 3.2 1.3 0.2] [4.6 3.1 1.5 0.2] [5. 3.6 1.4 0.2]]
First 5 labels of y (Target):[0 0 0 0 0]
Shape of X: (150, 4)Shape of y: (150,)为了评估模型在未见过数据上的性能,将数据集划分为训练集 (training set) 和测试集 (testing set) 至关重要。模型从训练集中学习,然后在测试集上评估其性能。这有助于防止过拟合 (overfitting),即模型在训练数据上表现良好,但在新数据上表现较差。
示例:训练集-测试集划分
Section titled “示例:训练集-测试集划分”sklearn.model_selection 中的 train_test_split 函数是常用的方法。这里,我们划分 Iris 数据集,分配 30% 用于测试,70% 用于训练。
from sklearn.datasets import load_irisfrom sklearn.model_selection import train_test_split
# Load dataX, y = load_iris(return_X_y=True)
# Split data: 70% training, 30% testingX_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, random_state=42, stratify=y)
print(f"Shape of X_train: {X_train.shape}")print(f"Shape of X_test: {X_test.shape}")print(f"Shape of y_train: {y_train.shape}")print(f"Shape of y_test: {y_test.shape}")Shape of X_train: (105, 4)Shape of X_test: (45, 4)Shape of y_train: (105,)Shape of y_test: (45,)train_test_split 的关键参数:
- X, y: 要划分的特征矩阵和响应向量。
- test_size: 测试集的比例大小(例如,0.3 表示 30%)。也可以使用
train_size。 - random_state: 用于随机数生成器的种子。设置此参数可确保每次运行代码时划分结果相同,这对于可重现性至关重要。可以使用任何整数值。
- stratify=y:(建议用于分类任务)确保训练集和测试集中目标类别比例与原始数据集相似。这对于不平衡数据集很重要。
准备好数据后,下一步是训练一个机器学习模型。Scikit-learn 提供了多种算法,它们都遵循一致的 API:fit() 用于训练,predict() 用于进行预测,以及 score() 用于评估。
示例:训练 K 近邻 (KNN) 分类器
Section titled “示例:训练 K 近邻 (KNN) 分类器”我们将使用 K 近邻 (KNN) 算法,这是一个简单但有效的分类器。此示例演示了基本的训练和预测工作流程。
from sklearn.datasets import load_irisfrom sklearn.model_selection import train_test_splitfrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.metrics import accuracy_scoreimport numpy as np
# 加载并划分数据(如前所示)X, y = load_iris(return_X_y=True)X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42, stratify=y)
# 初始化 KNN 分类器# n_neighbors 指定了 KNN 中的 'K'knn_classifier = KNeighborsClassifier(n_neighbors=3)
# 使用训练数据训练模型knn_classifier.fit(X_train, y_train)
# 在测试数据上进行预测y_pred = knn_classifier.predict(X_test)
# 评估模型的准确率accuracy = accuracy_score(y_test, y_pred)print(f"Accuracy on test set: {accuracy:.4f}")
# 对新的、未见过的数据样本进行预测示例# 注意:新的样本应具有与训练数据相同的特征数量(本例中为 4)sample_new_data = np.array([ [5.0, 3.1, 1.6, 0.2], # 期望结果:setosa (0) [6.5, 3.0, 5.2, 2.0] # 期望结果:virginica (2)])predicted_species_indices = knn_classifier.predict(sample_new_data)iris = load_iris() # 重新加载以便获取 target_namespredicted_species_names = iris.target_names[predicted_species_indices]print(f"Predictions for new samples: {predicted_species_names}")Accuracy on test set: 0.9778Predictions for new samples: ['setosa' 'virginica']关于 KNN 的更多细节,请参阅其专门的章节。选择正确的模型和调整其参数(例如 KNN 的 n_neighbors)通常涉及交叉验证 (cross-validation) 等技术,这将在后面讨论。
训练模型可能很耗时。模型训练完成后,您通常希望保存(持久化,persist)它以便将来使用,而无需重新训练。joblib 库(SciPy 生态系统的一部分,并由 Scikit-learn 推荐)非常适合此目的,特别是对于包含大型 NumPy 数组的对象。
让我们保存之前训练好的 knn_classifier 模型:
import joblib
# 假设 knn_classifier 是我们上一步训练好的模型# 将模型保存到文件model_filename = 'iris_knn_classifier.joblib'joblib.dump(knn_classifier, model_filename)print(f"Model saved to {model_filename}")
# ... 稍后,在另一个脚本或会话中 ...
# 从文件加载模型loaded_knn_model = joblib.load(model_filename)print(f"Model loaded from {model_filename}")
# 现在可以使用加载的模型 loaded_knn_model 进行预测# 示例:y_pred_loaded = loaded_knn_model.predict(X_test)# print(f"加载模型的准确率: {accuracy_score(y_test, y_pred_loaded):.4f}")这将保存整个模型对象,包括其学习到的参数。请确保加载模型的环境中具有相同版本的相关库(如 Scikit-learn),以保证兼容性。
原始数据通常不适合直接输入到机器学习算法中。预处理 (Preprocessing) 将原始数据转换为更可用的格式。Scikit-learn 的 sklearn.preprocessing 模块提供了各种工具。常见的预处理步骤包括缩放 (scaling)、归一化 (normalization) 和分类特征编码 (encoding categorical features)。
为什么需要预处理?许多算法(例如,SVM、KNN、PCA、神经网络)对特征尺度 (feature scales) 很敏感。具有较大值的特征可能会主导学习过程。预处理有助于标准化数据。
二值化 (Binarization) 根据阈值将数值特征转换为布尔值(0 或 1)。这对于从连续特征创建二元特征很有用。
import numpy as npfrom sklearn.preprocessing import Binarizer
input_data = np.array([ [2.1, -1.9, 5.5], [-1.5, 2.4, 3.5], [0.5, -7.9, 5.6], [5.9, 2.3, -5.8]])
# 大于 0.5 的值变为 1,否则变为 0binarizer = Binarizer(threshold=0.5)data_binarized = binarizer.fit_transform(input_data)print(f"\nBinarized data (threshold=0.5):\n{data_binarized}")Binarized data (threshold=0.5):[[1. 0. 1.] [0. 1. 1.] [0. 0. 1.] [1. 1. 0.]]去除均值(标准化)
Section titled “去除均值(标准化)”去除均值,更常见地称为标准化 (standardization) 或 Z-score 归一化 (Z-score normalization),将特征转换为具有零均值和单位方差。这是通过减去均值并除以每个特征的标准差来实现的。
import numpy as npfrom sklearn.preprocessing import StandardScaler
input_data = np.array([ [2.1, -1.9, 5.5], [-1.5, 2.4, 3.5], [0.5, -7.9, 5.6], [5.9, 2.3, -5.8]], dtype=np.float64)
print(f"Original Mean: {input_data.mean(axis=0)}")print(f"Original Std Dev: {input_data.std(axis=0)}")
scaler = StandardScaler()data_scaled = scaler.fit_transform(input_data)
print(f"\nScaled Data (StandardScaler):\n{data_scaled}")print(f"Mean of scaled data: {data_scaled.mean(axis=0)}")print(f"Std Dev of scaled data: {data_scaled.std(axis=0)}")Original Mean: [ 1.75 -1.275 2.2 ]Original Std Dev: [2.71431391 4.20022321 4.69414529]
Scaled Data (StandardScaler):[[ 0.12896806 0.6318175 0.70299018] [-1.19734452 0.87470814 0.2769086 ] [-0.46049308 -1.57685687 0.72410327] [ 1.52886954 0.07033123 -1.7039976 ]]Mean of scaled data: [2.22044605e-17 0.00000000e+00 0.00000000e+00]Std Dev of scaled data: [1. 1. 1.]注意:关键在于只在训练数据上 fit 缩放器 (scaler),然后在训练数据和测试数据上都使用 transform,以防止测试集中的数据泄露 (data leakage)。
缩放(最小-最大缩放)
Section titled “缩放(最小-最大缩放)”最小-最大缩放 (Min-Max scaling)(或归一化,normalization)将特征重新缩放到特定范围,通常是 [0, 1] 或 [-1, 1]。这是通过减去最小值并除以每个特征的范围(最大值 - 最小值)来完成的。
import numpy as npfrom sklearn.preprocessing import MinMaxScaler
input_data = np.array([ [2.1, -1.9, 5.5], [-1.5, 2.4, 3.5], [0.5, -7.9, 5.6], [5.9, 2.3, -5.8]], dtype=np.float64)
# 将特征缩放到 [0, 1] 范围min_max_scaler = MinMaxScaler(feature_range=(0, 1))data_scaled_minmax = min_max_scaler.fit_transform(input_data)print(f"\nMin-Max Scaled Data (range [0, 1]):\n{data_scaled_minmax}")Min-Max Scaled Data (range [0, 1]):[[0.48648649 0.58252427 0.99122807] [0. 1. 0.81578947] [0.27027027 0. 1. ] [1. 0.99029126 0. ]]Scikit-learn 中的归一化 (Normalization)(使用 sklearn.preprocessing.Normalizer)独立地重新缩放每个样本(行),使其具有单位范数 (unit norm)(L1 或 L2)。这与 StandardScaler 或 MinMaxScaler 等按特征(列)进行缩放不同。归一化通常用于文本分类或聚类。
L1 归一化(最小绝对误差)
Section titled “L1 归一化(最小绝对误差)”重新缩放每个样本,使其特征绝对值之和为 1。
import numpy as npfrom sklearn.preprocessing import Normalizer
input_data = np.array([ [2.1, -1.9, 5.5], [-1.5, 2.4, 3.5], [0.5, -7.9, 5.6], [5.9, 2.3, -5.8]], dtype=np.float64)
normalizer_l1 = Normalizer(norm='l1')data_normalized_l1 = normalizer_l1.fit_transform(input_data)print(f"\nL1 Normalized Data:\n{data_normalized_l1}")print(f"Sum of absolute values per row (L1 norm): {np.sum(np.abs(data_normalized_l1), axis=1)}")L1 Normalized Data:[[ 0.22105263 -0.2 0.57894737] [-0.2027027 0.32432432 0.47297297] [ 0.03571429 -0.56428571 0.4 ] [ 0.42142857 0.16428571 -0.41428571]]Sum of absolute values per row (L1 norm): [1. 1. 1. 1.]L2 归一化(最小二乘)
Section titled “L2 归一化(最小二乘)”重新缩放每个样本,使其特征平方和(欧几里得范数,Euclidean norm)为 1。
import numpy as npfrom sklearn.preprocessing import Normalizer
input_data = np.array([ [2.1, -1.9, 5.5], [-1.5, 2.4, 3.5], [0.5, -7.9, 5.6], [5.9, 2.3, -5.8]], dtype=np.float64)
normalizer_l2 = Normalizer(norm='l2')data_normalized_l2 = normalizer_l2.fit_transform(input_data)print(f"\nL2 Normalized Data:\n{data_normalized_l2}")print(f"Sum of squares per row (L2 norm squared): {np.sum(data_normalized_l2**2, axis=1)}")L2 Normalized Data:[[ 0.33946114 -0.30713151 0.88906489] [-0.33325106 0.53320169 0.7775858 ] [ 0.05156558 -0.81473612 0.57753446] [ 0.68706914 0.26784051 -0.6754239 ]]Sum of squares per row (L2 norm squared): [1. 1. 1. 1.]使用 Pipeline 简化工作流程
Section titled “使用 Pipeline 简化工作流程”Scikit-learn 的 Pipeline 对象允许您将多个处理步骤(如缩放和 PCA)与最终的估计器 (estimator)(如分类器,classifier)串联起来。强烈建议使用它来实现:
- 便捷性:将多个步骤组合成一个对象。
- 防止数据泄露:确保在交叉验证期间仅在训练数据上拟合 (fit) 变换。
- 可重现性:封装整个预处理和模型构建工作流程。
简单 Pipeline 示例:
from sklearn.pipeline import Pipelinefrom sklearn.preprocessing import StandardScalerfrom sklearn.svm import SVCfrom sklearn.datasets import load_irisfrom sklearn.model_selection import train_test_splitfrom sklearn.metrics import accuracy_score
# 加载并划分数据X, y = load_iris(return_X_y=True)X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# 创建 Pipeline:StandardScaler 然后 SVCpipe = Pipeline([ ('scaler', StandardScaler()), # 步骤 1:缩放数据 ('svc', SVC()) # 步骤 2:支持向量分类器])
# 拟合 Pipeline(应用 scaler 的 fit_transform,然后应用 SVC 的 fit)pipe.fit(X_train, y_train)
# 进行预测(应用 scaler 的 transform,然后应用 SVC 的 predict)y_pred_pipeline = pipe.predict(X_test)
accuracy_pipeline = accuracy_score(y_test, y_pred_pipeline)print(f"Accuracy with Pipeline: {accuracy_pipeline:.4f}")Accuracy with Pipeline: 1.0000Pipeline 是构建健壮且可维护的机器学习模型的强大工具。您可以在 Scikit-learn 关于 sklearn.pipeline.Pipeline 的文档中了解更多。