Skip to content

Scikit Learn - 扩展线性建模

Scikit-Learn - 使用多项式特征和 Pipeline 扩展线性模型

Section titled “Scikit-Learn - 使用多项式特征和 Pipeline 扩展线性模型”

本章重点介绍如何通过引入非线性特征(特别是多项式特征)来增强线性模型的性能,并使用 Scikit-Learn 的 Pipeline 工具来简化工作流程。

多项式特征 (Polynomial Features) 简介

Section titled “多项式特征 (Polynomial Features) 简介”

顾名思义,线性模型假设特征和目标变量之间存在线性关系。然而,现实世界的数据常常表现出非线性模式。允许线性模型拟合非线性数据的一种常用方法是将输入特征转换为更高维度的空间,该空间包含原始特征的非线性组合。多项式特征是对此的一种流行选择。

例如,对于一维数据的简单线性回归模型是 y = w_0 + w_1*x。如果关系是二次的,我们可以创建一个新特征 x^2 并拟合一个像 y = w_0 + w_1*x + w_2*x^2 这样的模型。这在系数 w_0, w_1, w_2 方面仍然是一个线性模型,但它可以捕捉与 x 的二次关系。

在数学上,对于二维数据 (x_1, x_2),标准的线性模型是:

ŷ(w, x) = w_0 + w_1*x_1 + w_2*x_2

通过构建二次多项式特征,包括交互项,模型变为:

ŷ(w, x) = w_0 + w_1*x_1 + w_2*x_2 + w_3*x_1*x_2 + w_4*x_1^2 + w_5*x_2^2

这个变换后的模型相对于新的特征集 (1, x_1, x_2, x_1*x_2, x_1^2, x_2^2) 保持线性,可以使用标准的线性模型技术进行求解。

Scikit-Learn 在 sklearn.preprocessing 中提供了 PolynomialFeatures 变换器来生成这些特征。

参数 (Parameter)描述 (Description)
degreeint 类型,默认值为 2。多项式特征的次数。例如,degree=2 生成的特征包含高达 x_i*x_j 和 x_i^2。
interaction_onlybool 类型,默认值为 False。如果为 True,则只生成交互特征:即至多由 degree 个不同的输入特征相乘构成的特征(例如,x_1*x_2 而非 x_1^2,如果 degree=2)。
include_biasbool 类型,默认值为 True。如果为 True,则包含一个偏置列,即所有多项式幂均为零的特征(常数项,1)。
order{‘C’, ‘F’},默认值为 ‘C’。密集情况下输出数组的顺序。‘F’ 顺序计算速度更快,但如果后续的估计器期望 C 连续数组,可能会减慢速度。
属性 (Attribute)描述 (Description)
powers_形状为 (n_output_features, n_input_features) 的数组。powers_[i, j] 是第 i 个输出特征中第 j 个输入特征的指数。
n_features_in_int 类型。在 fit 期间看到的特征数量(原始输入特征的数量)。
n_output_features_int 类型。多项式输出特征的总数。这也是变换后数据中的列数。

将一个二维数组变换为 degree 为 2 的多项式特征:

from sklearn.preprocessing import PolynomialFeatures
import numpy as np
# 样本输入数据:4 个样本,2 个特征
X_original = np.arange(8).reshape(4, 2)
# X_original 将是:
# [[0 1]
# [2 3]
# [4 5]
# [6 7]]
print(f"Original X:\n{X_original}")
# 初始化 PolynomialFeatures 变换器, degree 为 2
poly = PolynomialFeatures(degree=2, include_bias=True)
# 拟合并变换数据
X_poly = poly.fit_transform(X_original)
print(f"\nPolynomial features (degree 2):\n{X_poly}")
print(f"\nFeature names (if available, or use get_feature_names_out()): {poly.get_feature_names_out(['x1', 'x2'])}")
print(f"\nPowers matrix:\n{poly.powers_}")
Original X:
[[0 1]
[2 3]
[4 5]
[6 7]]
Polynomial features (degree 2):
[[ 1. 0. 1. 0. 0. 1.]
[ 1. 2. 3. 4. 6. 9.]
[ 1. 4. 5. 16. 20. 25.]
[ 1. 6. 7. 36. 42. 49.]]
Feature names (if available, or use get_feature_names_out()): ['1' 'x1' 'x2' 'x1^2' 'x1 x2' 'x2^2']
Powers matrix:
[[0 0]
[1 0]
[0 1]
[2 0]
[1 1]
[0 2]]

输出的特征是 [1, x1, x2, x1^2, x1*x2, x2^2]。

使用 Pipeline 工具 (sklearn.pipeline.Pipeline) 简化流程

Section titled “使用 Pipeline 工具 (sklearn.pipeline.Pipeline) 简化流程”

生成多项式特征等预处理步骤通常与拟合最终估计器相结合。Scikit-Learn 的 Pipeline 对象允许将多个估计器(变换器和最终预测器)链接成一个单一的元估计器。这简化了工作流程,防止了数据从测试集泄漏到训练过程中(例如,只在训练数据上拟合变换器),并使得对所有步骤的参数进行网格搜索更加容易。

示例:使用 Pipeline 进行多项式回归

Section titled “示例:使用 Pipeline 进行多项式回归”

让我们创建一个先生成多项式特征,然后拟合线性回归模型的 Pipeline。我们将尝试恢复已知多项式函数的系数。

from sklearn.linear_model import LinearRegression
from sklearn.pipeline import Pipeline # 或 make_pipeline 用于更简单的构建
# (PolynomialFeatures, np 已导入)
# 定义真实的多项式关系: y = 3 - 2*x + x^2 - x^3
np.random.seed(0)
x_true = np.arange(5)
y_true = 3 - 2 * x_true + x_true**2 - x_true**3
# 为 Scikit-Learn 重塑 x (作为列向量)
X_true_reshaped = x_true[:, np.newaxis]
# 创建一个 Pipeline:1. PolynomialFeatures,2. LinearRegression
# 我们需要 degree 为 3 的多项式特征。
# 对于 LinearRegression,fit_intercept=False 因为 PolynomialFeatures (include_bias=True) 已经提供了截距项。
poly_regression_pipeline = Pipeline([
('poly_features', PolynomialFeatures(degree=3, include_bias=True)),
('linear_regression', LinearRegression(fit_intercept=False))
])
# 拟合数据到 Pipeline
poly_regression_pipeline.fit(X_true_reshaped, y_true)
# 访问 Pipeline 中线性回归模型的系数
# 这些系数对应于 PolynomialFeatures 生成的 [1, x, x^2, x^3] 特征
learned_coefficients = poly_regression_pipeline.named_steps['linear_regression'].coef_
print(f"Learned coefficients: {learned_coefficients}")
# 在无噪声数据上的完美拟合,这些系数应该接近 [3., -2., 1., -1.]
# 使用 Pipeline 进行预测
# y_pred = poly_regression_pipeline.predict(X_true_reshaped)
# print(f"原始 y: {y_true}")
# print(f"预测 y: {np.round(y_pred, 2)}")
Learned coefficients: [ 3. -2. 1. -1.]

输出显示,在多项式特征上训练的线性模型(在 Pipeline 中)成功地恢复了输入多项式函数的精确系数。这展示了将特征工程与线性模型结合使用的强大功能以及使用 Pipeline 的便利性。

使用 Pipeline 的好处:

  • **方便性和封装性:**将多个处理步骤组合到一个对象中。
  • **联合参数选择:**允许同时对 Pipeline 中所有步骤的参数进行网格搜索。
  • **避免数据泄露:**确保预处理步骤(如缩放或 PCA)的拟合仅在交叉验证期间发生在训练数据上。

使用多项式特征可能会增加模型的复杂性并带来过拟合的风险,尤其是在次数很高的情况下。正则化(例如,在 Pipeline 中使用 Ridge 或 Lasso 而非 LinearRegression)通常是必要的。有关更多信息,请参阅 PolynomialFeatures 文档 和 Pipeline 文档。