Skip to content

随机梯度下降

本章介绍随机梯度下降(Stochastic Gradient Descent,简称 SGD),这是 Scikit-Learn 中一种强大且广泛使用的优化算法,用于拟合线性模型。

随机梯度下降(SGD)是一种高效的优化算法,用于寻找能最小化成本函数(或损失函数)的函数参数(系数)。它特别适用于大规模数据集,因为它每次只处理一个训练实例(或一小批量)来更新模型的系数,而不是像其他方法那样每次更新都处理整个数据集。这使得它具有可扩展性,适用于在线学习。SGD 常用于判别学习,为各种(通常是凸的)损失函数下的线性分类器和回归器训练模型,例如支持向量机(Support Vector Machines,简称 SVM)和逻辑回归(Logistic Regression)。

Scikit-Learn 中的 SGDClassifier 实现了基本的 SGD 学习过程。它支持多种损失函数和惩罚项,使其能够拟合不同的线性分类模型。

sklearn.linear_model.SGDClassifier 的重要参数:

ParameterDescription
lossstr 类型,默认为 ‘hinge’。要使用的损失函数。常用选项包括:‘hinge’(线性支持向量机)、‘log_loss’(逻辑回归)、‘modified_huber’(平滑损失,对离群值不敏感,产生概率估计)、‘squared_hinge’(二次惩罚合页损失)、‘perceptron’(感知机)。还有其他一些专用的损失函数。
penalty{‘l2’, ‘l1’, ‘elasticnet’, None},默认为 ‘l2’。正则化(惩罚)项。‘l2’ 是标准的 Ridge 正则化,‘l1’ 是 Lasso(可以产生稀疏解),‘elasticnet’ 结合了 L1 和 L2 正则化。None 表示不使用正则化。
alphafloat 类型,默认为 0.0001。乘以正则化项的常数。值越大表示正则化强度越大。
l1_ratiofloat 类型,默认为 0.15。Elastic Net 的混合参数,取值范围为 0 <= l1_ratio <= 1。l1_ratio=0 对应 L2 惩罚项,l1_ratio=1 对应 L1 惩罚项。仅在 penalty='elasticnet' 时使用。
fit_interceptbool 类型,默认为 True。是否估算截距。如果为 False,则假定数据已经中心化。
max_iterint 类型,默认为 1000。在训练数据上的最大遍历次数(epochs,轮次)。
tolfloat 或 None 类型,默认为 1e-3。停止训练的准则。如果不为 None,当连续 n_iter_no_change 个 epoch 满足 (loss > best_loss - tol) 时,训练将停止。
shufflebool 类型,默认为 True。每个 epoch 后是否打乱(洗牌)训练数据。
epsilonfloat 类型,默认为 0.1。ε-不敏感损失函数中的 ε(例如 ‘huber’、‘epsilon_insensitive’——尽管这些更常用于回归器)。对于 ‘modified_huber’,它是损失函数变为线性的点。
n_jobsint 类型,默认为 None。用于 One-Vs-Rest (‘ovr’) 计算的 CPU 核心数。None 表示 1 个,-1 表示使用所有 CPU。这仅用于多类别分类。
random_stateint, RandomState 实例或 None 类型,默认为 None。用于打乱数据时的随机种子。
learning_ratestr 类型,默认为 ‘optimal’。学习率调度策略:‘constant’:eta = eta0;‘optimal’:eta = 1.0 / (alpha * (t + t0))(其中 t0 是启发式值);‘invscaling’:eta = eta0 / pow(t, power_t);‘adaptive’:eta = eta0,只要训练损失持续下降。如果 early_stopping=True,每当连续 n_iter_no_change 个 epoch 训练损失未能下降或验证分数未能提高时,eta0 将减小 5 倍。
eta0double 类型,默认为 0.0。用于 ‘constant’、‘invscaling’、‘adaptive’ 策略的初始学习率。对于 ‘adaptive’ 默认为 0.01,否则为 0.0。
power_tdouble 类型,默认为 0.5。逆比例缩放学习率('invscaling')的指数。
early_stoppingbool 类型,默认为 False。是否使用早期停止策略,在验证分数没有提高时终止训练。如果为 True,会自动将一部分训练数据划分为验证集。
validation_fractionfloat 类型,默认为 0.1。用于早期停止策略时,划分为验证集的训练数据比例。仅在 early_stopping 为 True 时使用。
n_iter_no_changeint 类型,默认为 5。在停止拟合之前,等待没有改进的迭代次数。当连续 n_iter_no_change 次迭代中 loss 或 score 的改进小于 tol 时,算法认为收敛。
class_weightdict 类型,格式为 {类别标签: 权重},也可为 ‘balanced’ 或 None,默认为 None。与类别关联的权重。如果为 ‘balanced’,权重与类别的频率成反比。
warm_startbool 类型,默认为 False。如果设置为 True,则重用上次调用 fit 的结果作为初始化,否则会清除之前的结果。
averagebool 或 int 类型,默认为 False。如果设置为 True,则计算平均 SGD 权重,并将结果存储在 coef_ 和 intercept_ 属性中。如果设置为大于 1 的整数,则在见到的样本总数达到 average 时开始平均计算。

拟合 SGDClassifier 后:

AttributeDescription
coef_array 类型,当 n_classes==2 时 shape 为 (1, n_features),否则为 (n_classes, n_features)。分配给特征的权重。
intercept_array 类型,当 n_classes==2 时 shape 为 (1,),否则为 (n_classes,)。决策函数中的独立项(偏置)。
n_iter_int 类型。达到停止准则前的实际迭代次数。
loss_function_估算器使用的具体 LossFunction 对象。

实现示例:

SGDClassifier 需要使用两个数组进行拟合:X(训练样本,shape 为 [n_samples, n_features])和 Y(目标值/类别标签,shape 为 [n_samples])。

import numpy as np
from sklearn.linear_model import SGDClassifier
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
# Sample data
X = np.array([[-1, -1], [-2, -1], [1, 1], [2, 1], [-1, 2], [2, -2]])
Y = np.array([1, 1, 2, 2, 1, 2]) # Two classes: 1 and 2
# It's highly recommended to scale data before using SGD
# Create a pipeline to scale data and then fit SGDClassifier
sgd_clf = make_pipeline(StandardScaler(),
SGDClassifier(loss='hinge', penalty='l2',
max_iter=1000, tol=1e-3, random_state=42))
sgd_clf.fit(X, Y)
print("Pipeline with SGDClassifier fitted.")
# To access the classifier itself from the pipeline:
# final_estimator = sgd_clf.named_steps['sgdclassifier']
# print(f"Coefficients: {final_estimator.coef_}")
# print(f"Intercept: {final_estimator.intercept_}")

print 语句的输出:

Pipeline with SGDClassifier fitted.

预测新值:

new_data_point = np.array([[2., 2.]])
prediction = sgd_clf.predict(new_data_point)
print(f"Prediction for {new_data_point}: {prediction}")

输出:

Prediction for [[2. 2.]]: [2]

访问系数和截距(从 pipeline 中的分类器):

final_estimator = sgd_clf.named_steps['sgdclassifier']
print(f"Coefficients (weights): {final_estimator.coef_}")
print(f"Intercept: {final_estimator.intercept_}")

输出(示例值,会因 random_state 和具体缩放情况而有所不同):

Coefficients (weights): [[ 8.60414862 -0.97295213]]
Intercept: [1.72006884]

使用 decision_function 计算到超平面的有符号距离:

decision_value = sgd_clf.decision_function(new_data_point)
print(f"Decision function value for {new_data_point}: {decision_value}")

输出(示例值):

Decision function value for [[2. 2.]]: [12.8626418]

SGDRegressor 为线性回归模型实现了 SGD,支持多种损失函数和惩罚项。

参数与 SGDClassifier 大体相似。主要区别在于:

  • loss: str 类型,默认为 ‘squared_error’。选项包括:‘squared_error’(普通最小二乘)、‘huber’(对离群值不敏感,优于 squared_error)、‘epsilon_insensitive’(忽略特定 ε 范围内的误差,用于支持向量回归 SVR)、‘squared_epsilon_insensitive’。
  • epsilon: float 类型,默认为 0.1。用于 ‘huber’、‘epsilon_insensitive’ 和 ‘squared_epsilon_insensitive’ 损失函数的 ε。定义了误差开始被惩罚的阈值。
  • power_t: double 类型,默认为 0.25(与 SGDClassifier 的 0.5 不同)。用于 ‘invscaling’ 学习率的指数。
  • 没有 class_weight 参数,因为它用于回归任务。

类似于 SGDClassifier。如果 average=True 或一个 int,coef_ 和 intercept_ 将存储平均值。t_ 是一个附加属性:

  • t_: int 类型。训练期间执行的权重更新次数。如果每次迭代都使用所有样本,则等于 n_iter_ * n_samples_。

实现示例:

import numpy as np
from sklearn.linear_model import SGDRegressor
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.datasets import make_regression # For generating regression data
# Generate sample regression data
X_reg, y_reg = make_regression(n_samples=100, n_features=5, random_state=0, noise=0.1)
# Create a pipeline with StandardScaler and SGDRegressor
sgd_reg = make_pipeline(StandardScaler(),
SGDRegressor(loss='squared_error', penalty='elasticnet', l1_ratio=0.2,
max_iter=1000, tol=1e-3, random_state=0, average=False))
sgd_reg.fit(X_reg, y_reg)
print("Pipeline with SGDRegressor fitted.")
# Access the regressor from the pipeline
final_regressor = sgd_reg.named_steps['sgdregressor']
print(f"Coefficients: {final_regressor.coef_}")
print(f"Intercept: {final_regressor.intercept_}")
print(f"Number of iterations: {final_regressor.n_iter_}")
# print(f"Total weight updates (t_): {final_regressor.t_}") # t_ might not be directly exposed in all versions or under all conditions in pipeline

输出(示例值):

Pipeline with SGDRegressor fitted.
Coefficients: [32.2253343 16.72874275 41.86994217 60.73392231 28.50582354]
Intercept: [0.02206273]
Number of iterations: 11

使用拟合好的回归器进行预测:

new_reg_data_point = X_reg[0].reshape(1, -1) # Take the first sample for prediction
reg_prediction = sgd_reg.predict(new_reg_data_point)
print(f"Prediction for first sample: {reg_prediction}, Actual value: {y_reg[0]}")

输出(示例值):

Prediction for first sample: [-7.8509133], Actual value: -7.915208493840426

优点:

  • 效率高:SGD 计算效率高,尤其适用于大规模数据集。
  • 可扩展性好:能有效地处理大规模问题(大量样本和/或特征)。
  • 支持在线学习:可以使用 partial_fit 方法随着新数据的到来进行更新。

缺点:

  • 对特征缩放敏感:SGD 对特征缩放非常敏感。在训练之前对特征进行缩放(例如使用 StandardScaler)至关重要。
  • 需要调优超参数:需要仔细调优多个超参数,例如学习率和正则化参数。
  • 收敛性:由于更新的随机性,通往最小值的路径可能会有噪声(不稳定)。

有关更多详细信息和高级用法,请参阅 Scikit-Learn SGD 文档。