Skip to content

Scikit Learn - Boosting 方法

Boosting 方法是强大的集成学习 (ensemble learning) 技术,通过依次训练多个弱学习器 (weak learners)(通常是浅层决策树 shallow decision trees)来构建一个强模型 (strong model)。后续的每个学习器都更侧重于之前学习器错误分类 (misclassified) 的实例,从而逐步改进 (incrementally improving) 整体模型。Scikit-learn 的 sklearn.ensemble 模块提供了几种流行的 boosting 算法。

AdaBoost (Adaptive Boosting, 自适应增强)

Section titled “AdaBoost (Adaptive Boosting, 自适应增强)”

AdaBoost 是最早且最成功的 boosting 算法之一。它通过在反复修改的数据版本上拟合一系列弱学习器来工作。然后通过加权多数投票 (weighted majority vote)(或求和)将所有弱学习器的预测结合起来,产生最终预测。每次 boosting 迭代的数据修改涉及对每个训练样本应用权重。最初,所有权重都设置为相等,但在每个步骤中,错误分类实例的权重会增加,以便后续的弱学习器更能关注这些困难的案例。

Scikit-learn 提供了 sklearn.ensemble.AdaBoostClassifier。关键参数是 base_estimator(或在新版本中是 estimator),它指定了弱学习器。如果设置为 None,则默认为 DecisionTreeClassifier(max_depth=1)(一个决策树桩 decision stump)。另一个重要参数是 n_estimators,即要训练的弱学习器数量。

from sklearn.ensemble import AdaBoostClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
# 生成合成分类数据集
X, y = make_classification(n_samples=1000, n_features=20, n_informative=5,
n_redundant=0, random_state=42, shuffle=False)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# 初始化 AdaBoostClassifier,使用默认基学习器 (DecisionTreeClassifier(max_depth=1))
adb_clf = AdaBoostClassifier(n_estimators=100, random_state=42, algorithm='SAMME.R') # SAMME.R 用于实数输出
adb_clf.fit(X_train, y_train)
y_pred = adb_clf.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
print(f"AdaBoostClassifier 准确率: {accuracy:.4f}")
# 使用不同基学习器的示例
# custom_base_estimator = DecisionTreeClassifier(max_depth=2)
# adb_clf_custom = AdaBoostClassifier(estimator=custom_base_estimator, n_estimators=50, random_state=42)
# adb_clf_custom.fit(X_train, y_train)
# ...
AdaBoostClassifier 准确率: 0.8600

示例:在乳腺癌数据集上使用 AdaBoostClassifier

Section titled “示例:在乳腺癌数据集上使用 AdaBoostClassifier”
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.ensemble import AdaBoostClassifier
import numpy as np
# 加载数据集
data = load_breast_cancer()
X, y = data.data, data.target
# 初始化 AdaBoostClassifier
adb_clf_cancer = AdaBoostClassifier(n_estimators=100, random_state=42)
# 使用交叉验证评估
# (此处没有显式训练/测试分割,cross_val_score 处理了它)
# 注意:为了进行可靠评估,如果基学习器对特征尺度敏感,可能需要考虑缩放特征
scores = cross_val_score(adb_clf_cancer, X, y, cv=5) # 5 折交叉验证
print(f"乳腺癌数据集上的 AdaBoost - 交叉验证得分: {scores}")
print(f"乳腺癌数据集上的 AdaBoost - 平均准确率: {np.mean(scores):.4f}")
乳腺癌数据集上的 AdaBoost - 交叉验证得分: [0.97368421 0.97368421 0.94736842 0.99122807 0.92920354]
乳腺癌数据集上的 AdaBoost - 平均准确率: 0.9630

Scikit-learn 提供了 sklearn.ensemble.AdaBoostRegressor。其工作原理与分类器类似,但弱学习器预测连续值,并且它们的组合旨在最小化回归损失 (regression loss)(例如,线性 linear、平方 square、指数 exponential)。

from sklearn.ensemble import AdaBoostRegressor
from sklearn.datasets import make_regression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
import numpy as np
# 生成合成回归数据集
X, y = make_regression(n_samples=1000, n_features=10, n_informative=5,
random_state=42, shuffle=False)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# 初始化 AdaBoostRegressor
adb_reg = AdaBoostRegressor(n_estimators=100, random_state=42, loss='square')
adb_reg.fit(X_train, y_train)
y_pred_reg = adb_reg.predict(X_test)
mse = mean_squared_error(y_test, y_pred_reg)
print(f"AdaBoostRegressor MSE: {mse:.4f}")
print(f"AdaBoostRegressor RMSE: {np.sqrt(mse):.4f}")
AdaBoostRegressor MSE: 6234.1421
AdaBoostRegressor RMSE: 78.9566

梯度树提升 (Gradient Tree Boosting, GBT / GBRT)

Section titled “梯度树提升 (Gradient Tree Boosting, GBT / GBRT)”

梯度树提升 (Gradient Tree Boosting,也称为 Gradient Boosted Regression Trees 或 GBRT) 是一种更通用的 boosting 算法。它以阶段式 (stage-wise) 的方式构建一个加法模型 (additive model);它允许优化任意可微的损失函数 (arbitrary differentiable loss functions)。在每个阶段,都在给定损失函数的负梯度 (negative gradient)(残差误差 residual errors)上拟合一个回归树 (regression tree)。GBT 高效 (highly effective) 且广泛应用 (widely used) 于分类和回归任务。

关键参数包括 n_estimators(boosting 阶段/树的数量)、learning_rate(缩小每棵树的贡献,有助于防止过拟合 overfitting)、max_depth(单个回归估计器 estimator 的最大深度)和 loss(要优化的损失函数)。

Scikit-learn 提供了 sklearn.ensemble.GradientBoostingClassifier。对于二元分类 (binary classification),常用的损失函数有 ‘deviance’(逻辑回归损失 logistic regression loss)或 ‘exponential’(类似于 AdaBoost)。对于多类别 (multi-class),使用 ‘deviance’(多类别逻辑损失 multi-class logistic loss)。

from sklearn.ensemble import GradientBoostingClassifier
from sklearn.datasets import make_hastie_10_2 # 常用于 boosting 示例的二元分类问题
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
# 生成数据集(make_hastie_10_2 常用于 boosting 示例)
X, y = make_hastie_10_2(random_state=0)
# 如果某些指标/损失函数需要,将 y 标签从 {-1, 1} 转换为 {0, 1}
# y = (y + 1) // 2
X_train, X_test = X[:5000], X[5000:]
y_train, y_test = y[:5000], y[5000:]
gbc_clf = GradientBoostingClassifier(n_estimators=100, learning_rate=0.1,
max_depth=3, random_state=0)
gbc_clf.fit(X_train, y_train)
y_pred_gbc = gbc_clf.predict(X_test)
accuracy_gbc = accuracy_score(y_test, y_pred_gbc)
# 对于 make_hastie_10_2,score() 直接工作,因为它期望 {-1, 1}
# 对于其他数据集,确保二元 'deviance' 的 y 是 0/1,或检查 predict_proba
# print(f"GradientBoostingClassifier 得分 (internal): {gbc_clf.score(X_test, y_test):.4f}")
print(f"GradientBoostingClassifier 准确率: {accuracy_gbc:.4f}")

输出(make_hastie_10_2 目标是 -1, 1;accuracy_score 会处理)

Section titled “输出(make_hastie_10_2 目标是 -1, 1;accuracy_score 会处理)”
GradientBoostingClassifier 准确率: 0.9143

示例:在乳腺癌数据集上使用 GradientBoostingClassifier

Section titled “示例:在乳腺癌数据集上使用 GradientBoostingClassifier”
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.ensemble import GradientBoostingClassifier
import numpy as np
# 加载数据集
data = load_breast_cancer()
X, y = data.data, data.target
gbc_cancer = GradientBoostingClassifier(n_estimators=100, learning_rate=0.1,
max_depth=3, random_state=42)
scores_gbc = cross_val_score(gbc_cancer, X, y, cv=5)
print(f"乳腺癌数据集上的 GradientBoosting - 交叉验证得分: {scores_gbc}")
print(f"乳腺癌数据集上的 GradientBoosting - 平均准确率: {np.mean(scores_gbc):.4f}")
乳腺癌数据集上的 GradientBoosting - 交叉验证得分: [0.96491228 0.96491228 0.95614035 0.98245614 0.94690265]
乳腺癌数据集上的 GradientBoosting - 平均准确率: 0.9631

Scikit-learn 提供了 sklearn.ensemble.GradientBoostingRegressor。常用的损失函数包括 ‘ls’(最小二乘 least squares)、‘lad’(最小绝对偏差 least absolute deviation)、‘huber’(ls 和 lad 的组合,对异常值 robust to outliers 鲁棒)和 ‘quantile’(分位数)。

from sklearn.ensemble import GradientBoostingRegressor
from sklearn.datasets import make_friedman1 # 具有非线性结构的回归问题
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
import numpy as np
X, y = make_friedman1(n_samples=2000, random_state=0, noise=1.0)
X_train, X_test = X[:1000], X[1000:]
y_train, y_test = y[:1000], y[1000:]
gbr_reg = GradientBoostingRegressor(n_estimators=100, learning_rate=0.1,
max_depth=3, random_state=0, loss='squared') # 'squared' 表示平方误差
gbr_reg.fit(X_train, y_train)
y_pred_gbr = gbr_reg.predict(X_test)
mse_gbr = mean_squared_error(y_test, y_pred_gbr)
print(f"GradientBoostingRegressor MSE: {mse_gbr:.4f}")
print(f"GradientBoostingRegressor RMSE: {np.sqrt(mse_gbr):.4f}")
GradientBoostingRegressor MSE: 5.0120
GradientBoostingRegressor RMSE: 2.2387

基于直方图的梯度提升 (Histogram-Based Gradient Boosting, HistGradientBoosting)

Section titled “基于直方图的梯度提升 (Histogram-Based Gradient Boosting, HistGradientBoosting)”

对于大型数据集,Scikit-learn 提供了 HistGradientBoostingClassifier 和 HistGradientBoostingRegressor。这些方法受到 LightGBM 的启发,比传统的 GradientBoostingClassifier/Regressor 显著更快 (significantly faster)。它们使用一种称为基于直方图的分箱 (histogram-based binning) 技术处理连续特征 (continuous features),这加快了树的构建速度 (speeds up tree construction) 并减少了内存使用 (reduces memory usage)。它们还内置支持处理缺失值 (handling missing values)。

from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=10000, n_features=20, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
hgb_clf = HistGradientBoostingClassifier(max_iter=100, learning_rate=0.1, random_state=42)
hgb_clf.fit(X_train, y_train)
y_pred_hgb = hgb_clf.predict(X_test)
accuracy_hgb = accuracy_score(y_test, y_pred_hgb)
print(f"HistGradientBoostingClassifier 准确率: {accuracy_hgb:.4f}")

输出(准确率可能因 scikit-learn 版本而略有不同)

Section titled “输出(准确率可能因 scikit-learn 版本而略有不同)”
HistGradientBoostingClassifier 准确率: 0.8987

Boosting 方法,尤其是 GBT 及其变体(如 XGBoost、LightGBM 和 CatBoost,它们是独立的库,但在生态系统中很流行),通常在机器学习竞赛和实际应用中表现顶尖 (top performers)。Scikit-learn 的实现提供了坚实的基础 (solid foundation),并且很好地集成在其 API 中 (well-integrated within its API)。

如需进一步探索 (further exploration),请查看 Scikit-learn 用户指南 (User Guide):https://scikit-learn.org/stable/modules/ensemble.html#gradient-boosting