随机梯度下降
Scikit-Learn - 随机梯度下降 (SGD)
Section titled “Scikit-Learn - 随机梯度下降 (SGD)”本章介绍随机梯度下降(Stochastic Gradient Descent,简称 SGD),这是 Scikit-Learn 中一种强大且广泛使用的优化算法,用于拟合线性模型。
随机梯度下降(SGD)是一种高效的优化算法,用于寻找能最小化成本函数(或损失函数)的函数参数(系数)。它特别适用于大规模数据集,因为它每次只处理一个训练实例(或一小批量)来更新模型的系数,而不是像其他方法那样每次更新都处理整个数据集。这使得它具有可扩展性,适用于在线学习。SGD 常用于判别学习,为各种(通常是凸的)损失函数下的线性分类器和回归器训练模型,例如支持向量机(Support Vector Machines,简称 SVM)和逻辑回归(Logistic Regression)。
SGDClassifier
Section titled “SGDClassifier”Scikit-Learn 中的 SGDClassifier 实现了基本的 SGD 学习过程。它支持多种损失函数和惩罚项,使其能够拟合不同的线性分类模型。
sklearn.linear_model.SGDClassifier 的重要参数:
| Parameter | Description |
|---|---|
| loss | str 类型,默认为 ‘hinge’。要使用的损失函数。常用选项包括:‘hinge’(线性支持向量机)、‘log_loss’(逻辑回归)、‘modified_huber’(平滑损失,对离群值不敏感,产生概率估计)、‘squared_hinge’(二次惩罚合页损失)、‘perceptron’(感知机)。还有其他一些专用的损失函数。 |
| penalty | {‘l2’, ‘l1’, ‘elasticnet’, None},默认为 ‘l2’。正则化(惩罚)项。‘l2’ 是标准的 Ridge 正则化,‘l1’ 是 Lasso(可以产生稀疏解),‘elasticnet’ 结合了 L1 和 L2 正则化。None 表示不使用正则化。 |
| alpha | float 类型,默认为 0.0001。乘以正则化项的常数。值越大表示正则化强度越大。 |
| l1_ratio | float 类型,默认为 0.15。Elastic Net 的混合参数,取值范围为 0 <= l1_ratio <= 1。l1_ratio=0 对应 L2 惩罚项,l1_ratio=1 对应 L1 惩罚项。仅在 penalty='elasticnet' 时使用。 |
| fit_intercept | bool 类型,默认为 True。是否估算截距。如果为 False,则假定数据已经中心化。 |
| max_iter | int 类型,默认为 1000。在训练数据上的最大遍历次数(epochs,轮次)。 |
| tol | float 或 None 类型,默认为 1e-3。停止训练的准则。如果不为 None,当连续 n_iter_no_change 个 epoch 满足 (loss > best_loss - tol) 时,训练将停止。 |
| shuffle | bool 类型,默认为 True。每个 epoch 后是否打乱(洗牌)训练数据。 |
| epsilon | float 类型,默认为 0.1。ε-不敏感损失函数中的 ε(例如 ‘huber’、‘epsilon_insensitive’——尽管这些更常用于回归器)。对于 ‘modified_huber’,它是损失函数变为线性的点。 |
| n_jobs | int 类型,默认为 None。用于 One-Vs-Rest (‘ovr’) 计算的 CPU 核心数。None 表示 1 个,-1 表示使用所有 CPU。这仅用于多类别分类。 |
| random_state | int, RandomState 实例或 None 类型,默认为 None。用于打乱数据时的随机种子。 |
| learning_rate | str 类型,默认为 ‘optimal’。学习率调度策略:‘constant’:eta = eta0;‘optimal’:eta = 1.0 / (alpha * (t + t0))(其中 t0 是启发式值);‘invscaling’:eta = eta0 / pow(t, power_t);‘adaptive’:eta = eta0,只要训练损失持续下降。如果 early_stopping=True,每当连续 n_iter_no_change 个 epoch 训练损失未能下降或验证分数未能提高时,eta0 将减小 5 倍。 |
| eta0 | double 类型,默认为 0.0。用于 ‘constant’、‘invscaling’、‘adaptive’ 策略的初始学习率。对于 ‘adaptive’ 默认为 0.01,否则为 0.0。 |
| power_t | double 类型,默认为 0.5。逆比例缩放学习率('invscaling')的指数。 |
| early_stopping | bool 类型,默认为 False。是否使用早期停止策略,在验证分数没有提高时终止训练。如果为 True,会自动将一部分训练数据划分为验证集。 |
| validation_fraction | float 类型,默认为 0.1。用于早期停止策略时,划分为验证集的训练数据比例。仅在 early_stopping 为 True 时使用。 |
| n_iter_no_change | int 类型,默认为 5。在停止拟合之前,等待没有改进的迭代次数。当连续 n_iter_no_change 次迭代中 loss 或 score 的改进小于 tol 时,算法认为收敛。 |
| class_weight | dict 类型,格式为 {类别标签: 权重},也可为 ‘balanced’ 或 None,默认为 None。与类别关联的权重。如果为 ‘balanced’,权重与类别的频率成反比。 |
| warm_start | bool 类型,默认为 False。如果设置为 True,则重用上次调用 fit 的结果作为初始化,否则会清除之前的结果。 |
| average | bool 或 int 类型,默认为 False。如果设置为 True,则计算平均 SGD 权重,并将结果存储在 coef_ 和 intercept_ 属性中。如果设置为大于 1 的整数,则在见到的样本总数达到 average 时开始平均计算。 |
拟合 SGDClassifier 后:
| Attribute | Description |
|---|---|
| coef_ | array 类型,当 n_classes==2 时 shape 为 (1, n_features),否则为 (n_classes, n_features)。分配给特征的权重。 |
| intercept_ | array 类型,当 n_classes==2 时 shape 为 (1,),否则为 (n_classes,)。决策函数中的独立项(偏置)。 |
| n_iter_ | int 类型。达到停止准则前的实际迭代次数。 |
| loss_function_ | 估算器使用的具体 LossFunction 对象。 |
实现示例:
SGDClassifier 需要使用两个数组进行拟合:X(训练样本,shape 为 [n_samples, n_features])和 Y(目标值/类别标签,shape 为 [n_samples])。
import numpy as npfrom sklearn.linear_model import SGDClassifierfrom sklearn.preprocessing import StandardScalerfrom sklearn.pipeline import make_pipeline
# Sample dataX = np.array([[-1, -1], [-2, -1], [1, 1], [2, 1], [-1, 2], [2, -2]])Y = np.array([1, 1, 2, 2, 1, 2]) # Two classes: 1 and 2
# It's highly recommended to scale data before using SGD# Create a pipeline to scale data and then fit SGDClassifiersgd_clf = make_pipeline(StandardScaler(), SGDClassifier(loss='hinge', penalty='l2', max_iter=1000, tol=1e-3, random_state=42))sgd_clf.fit(X, Y)print("Pipeline with SGDClassifier fitted.")# To access the classifier itself from the pipeline:# final_estimator = sgd_clf.named_steps['sgdclassifier']# print(f"Coefficients: {final_estimator.coef_}")# print(f"Intercept: {final_estimator.intercept_}")print 语句的输出:
Pipeline with SGDClassifier fitted.预测新值:
new_data_point = np.array([[2., 2.]])prediction = sgd_clf.predict(new_data_point)print(f"Prediction for {new_data_point}: {prediction}")输出:
Prediction for [[2. 2.]]: [2]访问系数和截距(从 pipeline 中的分类器):
final_estimator = sgd_clf.named_steps['sgdclassifier']print(f"Coefficients (weights): {final_estimator.coef_}")print(f"Intercept: {final_estimator.intercept_}")输出(示例值,会因 random_state 和具体缩放情况而有所不同):
Coefficients (weights): [[ 8.60414862 -0.97295213]]Intercept: [1.72006884]使用 decision_function 计算到超平面的有符号距离:
decision_value = sgd_clf.decision_function(new_data_point)print(f"Decision function value for {new_data_point}: {decision_value}")输出(示例值):
Decision function value for [[2. 2.]]: [12.8626418]SGDRegressor
Section titled “SGDRegressor”SGDRegressor 为线性回归模型实现了 SGD,支持多种损失函数和惩罚项。
参数与 SGDClassifier 大体相似。主要区别在于:
- loss: str 类型,默认为 ‘squared_error’。选项包括:‘squared_error’(普通最小二乘)、‘huber’(对离群值不敏感,优于 squared_error)、‘epsilon_insensitive’(忽略特定 ε 范围内的误差,用于支持向量回归 SVR)、‘squared_epsilon_insensitive’。
- epsilon: float 类型,默认为 0.1。用于 ‘huber’、‘epsilon_insensitive’ 和 ‘squared_epsilon_insensitive’ 损失函数的 ε。定义了误差开始被惩罚的阈值。
- power_t: double 类型,默认为 0.25(与
SGDClassifier的 0.5 不同)。用于 ‘invscaling’ 学习率的指数。 - 没有
class_weight参数,因为它用于回归任务。
类似于 SGDClassifier。如果 average=True 或一个 int,coef_ 和 intercept_ 将存储平均值。t_ 是一个附加属性:
- t_: int 类型。训练期间执行的权重更新次数。如果每次迭代都使用所有样本,则等于
n_iter_ * n_samples_。
实现示例:
import numpy as npfrom sklearn.linear_model import SGDRegressorfrom sklearn.preprocessing import StandardScalerfrom sklearn.pipeline import make_pipelinefrom sklearn.datasets import make_regression # For generating regression data
# Generate sample regression dataX_reg, y_reg = make_regression(n_samples=100, n_features=5, random_state=0, noise=0.1)
# Create a pipeline with StandardScaler and SGDRegressorsgd_reg = make_pipeline(StandardScaler(), SGDRegressor(loss='squared_error', penalty='elasticnet', l1_ratio=0.2, max_iter=1000, tol=1e-3, random_state=0, average=False))sgd_reg.fit(X_reg, y_reg)print("Pipeline with SGDRegressor fitted.")
# Access the regressor from the pipelinefinal_regressor = sgd_reg.named_steps['sgdregressor']print(f"Coefficients: {final_regressor.coef_}")print(f"Intercept: {final_regressor.intercept_}")print(f"Number of iterations: {final_regressor.n_iter_}")# print(f"Total weight updates (t_): {final_regressor.t_}") # t_ might not be directly exposed in all versions or under all conditions in pipeline输出(示例值):
Pipeline with SGDRegressor fitted.Coefficients: [32.2253343 16.72874275 41.86994217 60.73392231 28.50582354]Intercept: [0.02206273]Number of iterations: 11使用拟合好的回归器进行预测:
new_reg_data_point = X_reg[0].reshape(1, -1) # Take the first sample for predictionreg_prediction = sgd_reg.predict(new_reg_data_point)print(f"Prediction for first sample: {reg_prediction}, Actual value: {y_reg[0]}")输出(示例值):
Prediction for first sample: [-7.8509133], Actual value: -7.915208493840426SGD 的优点与缺点
Section titled “SGD 的优点与缺点”优点:
- 效率高:SGD 计算效率高,尤其适用于大规模数据集。
- 可扩展性好:能有效地处理大规模问题(大量样本和/或特征)。
- 支持在线学习:可以使用
partial_fit方法随着新数据的到来进行更新。
缺点:
- 对特征缩放敏感:SGD 对特征缩放非常敏感。在训练之前对特征进行缩放(例如使用
StandardScaler)至关重要。 - 需要调优超参数:需要仔细调优多个超参数,例如学习率和正则化参数。
- 收敛性:由于更新的随机性,通往最小值的路径可能会有噪声(不稳定)。
有关更多详细信息和高级用法,请参阅 Scikit-Learn SGD 文档。