Skip to content

随机决策树

本章探讨 Scikit-Learn 中的随机决策树算法,它们是强大的集成(ensemble)方法,以高准确性和抗过拟合能力而闻名。

标准决策树容易发生过拟合,特别是当树的深度较大时。随机决策树算法,如随机森林(Random Forests)和极度随机树(Extra-Trees),通过构建多个决策树的集成(ensemble)来解决这个问题。随机性主要通过两种方式引入:在数据的一个不同子样本(subsample)(装袋 bagging)上训练每棵树,和/或在每个分裂点只考虑特征的一个随机子集。最终的预测通常是单个树的预测结果的平均值(用于回归)或多数投票(用于分类)。

sklearn.ensemble 模块提供了这些算法的实现。

随机森林构建多个决策树。对于每棵树,它通常使用训练数据的自助采样(bootstrap sample,有放回的随机采样)。在树中分裂节点时,它只考虑特征的一个随机子集来寻找最佳分裂。这种双重随机性有助于降低树之间的相关性,减少方差并提高泛化能力。随机森林可用于分类和回归任务。

随机森林分类器 (RandomForestClassifier)

Section titled “随机森林分类器 (RandomForestClassifier)”

sklearn.ensemble.RandomForestClassifier 用于分类任务。关键参数包括 n_estimators(树的数量)、max_features(随机特征子集的大小)、max_depth(每棵树的最大深度)和 min_samples_split(分裂节点所需的最小样本数)。

from sklearn.model_selection import cross_val_score
from sklearn.datasets import make_blobs
from sklearn.ensemble import RandomForestClassifier
# Generate synthetic data for classification
X, y = make_blobs(n_samples=1000, n_features=10, centers=5, random_state=0)
# Initialize Random Forest Classifier
# Using fewer estimators for quicker example, in practice use more (e.g., 100+)
rf_clf = RandomForestClassifier(n_estimators=50, max_depth=None,
min_samples_split=2, random_state=0)
# Evaluate using cross-validation
scores = cross_val_score(rf_clf, X, y, cv=5) # 5-fold cross-validation
print(f"Cross-validation scores: {scores}")
print(f"Mean cross-validation score: {scores.mean():.4f}")
Cross-validation scores: [0.975 0.97 0.98 0.965 0.975]
Mean cross-validation score: 0.9730

使用 Iris 数据集进行一个实际的分类示例,包括评估指标。

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report, confusion_matrix, accuracy_score
from sklearn.datasets import load_iris
# Load Iris dataset
iris = load_iris()
X_iris = iris.data
y_iris = iris.target
# Split data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X_iris, y_iris, test_size=0.30, random_state=42)
# Initialize and train the Random Forest Classifier
rf_iris_clf = RandomForestClassifier(n_estimators=100, random_state=42)
rf_iris_clf.fit(X_train, y_train)
# Make predictions
y_pred = rf_iris_clf.predict(X_test)
# Evaluate the classifier
print("Confusion Matrix:")
print(confusion_matrix(y_test, y_pred))
print("\nClassification Report:")
print(classification_report(y_test, y_pred, target_names=iris.target_names))
print(f"Accuracy Score: {accuracy_score(y_test, y_pred):.4f}")
Confusion Matrix:
[[19 0 0]
[ 0 13 0]
[ 0 0 13]]
Classification Report:
precision recall f1-score support
setosa 1.00 1.00 1.00 19
versicolor 1.00 1.00 1.00 13
virginica 1.00 1.00 1.00 13
accuracy 1.00 45
macro avg 1.00 1.00 1.00 45
weighted avg 1.00 1.00 1.00 45
Accuracy Score: 1.0000

随机森林回归器 (RandomForestRegressor)

Section titled “随机森林回归器 (RandomForestRegressor)”

sklearn.ensemble.RandomForestRegressor 用于回归任务。它使用与分类器相似的参数,但分裂节点的准则通常基于减少均方误差(Mean Squared Error, MSE)或平均绝对误差(Mean Absolute Error, MAE)。

from sklearn.ensemble import RandomForestRegressor
from sklearn.datasets import make_regression
import numpy as np
# Generate synthetic data for regression
X_reg, y_reg = make_regression(n_features=10, n_informative=5, random_state=0, shuffle=False)
# Initialize and train Random Forest Regressor
rf_regr = RandomForestRegressor(n_estimators=100, max_depth=10, random_state=0)
rf_regr.fit(X_reg, y_reg)
# Make a prediction on a sample point (e.g., the first point from X_reg)
sample_point = X_reg[0].reshape(1, -1)
prediction = rf_regr.predict(sample_point)
print(f"Model parameters: {rf_regr.get_params()['n_estimators']} estimators, max_depth {rf_regr.get_params()['max_depth']}")
print(f"Prediction for sample point {sample_point}: {prediction}")
print(f"Actual value for sample point: {y_reg[0]}")
Model parameters: 100 estimators, max_depth 10
Prediction for sample point [[ 0.18500369 0.89837401 0.75398021 -0.40579343 0.01871914 -0.7394844
-0.84078684 -0.46938239 0.17930832 0.00304159]]: [-47.6763054]
Actual value for sample point: -45.04868412282407

与随机森林相比,极度随机树引入了更多的随机性。在选择分裂点时,不仅考虑特征的一个随机子集,而且对于每个考虑的特征,分裂阈值也是随机选择的(从当前样本中该特征范围内的均匀分布中选取)。这通常会导致偏差(bias)略高,但可以进一步减少方差(variance)。由于它们不需要搜索最优阈值,通常训练速度更快。

极度随机树分类器 (ExtraTreesClassifier)

Section titled “极度随机树分类器 (ExtraTreesClassifier)”

sklearn.ensemble.ExtraTreesClassifier 使用与 RandomForestClassifier 相同的关键参数。

from sklearn.model_selection import cross_val_score
from sklearn.datasets import make_blobs
from sklearn.ensemble import ExtraTreesClassifier
# Re-use synthetic data from Random Forest example
X_et, y_et = make_blobs(n_samples=1000, n_features=10, centers=5, random_state=0)
# Initialize Extra Trees Classifier
et_clf = ExtraTreesClassifier(n_estimators=50, max_depth=None,
min_samples_split=2, random_state=0)
# Evaluate using cross-validation
et_scores = cross_val_score(et_clf, X_et, y_et, cv=5)
print(f"Extra-Trees cross-validation scores: {et_scores}")
print(f"Extra-Trees mean cross-validation score: {et_scores.mean():.4f}")
Extra-Trees cross-validation scores: [0.98 0.97 0.98 0.97 0.98 ]
Extra-Trees mean cross-validation score: 0.9760

实现示例 2:使用不同的数据集(例如 make_classification)

Section titled “实现示例 2:使用不同的数据集(例如 make_classification)”

使用 make_classification 处理更复杂的场景。

from sklearn.datasets import make_classification
from sklearn.model_selection import KFold
# (KFold and cross_val_score already imported)
# (ExtraTreesClassifier already imported)
X_class, Y_class = make_classification(n_samples=1000, n_features=20, n_informative=15,
n_redundant=5, random_state=7, n_classes=3)
# Setup KFold for cross-validation
kfold = KFold(n_splits=10, random_state=7, shuffle=True) # Added shuffle=True for reproducibility with random_state
num_trees = 150
max_feat = 7 # Example: consider 7 features at each split
et_clf_complex = ExtraTreesClassifier(n_estimators=num_trees, max_features=max_feat, random_state=7)
results_complex = cross_val_score(et_clf_complex, X_class, Y_class, cv=kfold)
print(f"Extra-Trees (complex data) mean accuracy: {results_complex.mean():.4f}")
Extra-Trees (complex data) mean accuracy: 0.8050

极度随机树回归器 (ExtraTreesRegressor)

Section titled “极度随机树回归器 (ExtraTreesRegressor)”

sklearn.ensemble.ExtraTreesRegressor 用于回归任务,其参数与 ExtraTreesClassifier 类似。

from sklearn.ensemble import ExtraTreesRegressor
# (make_regression and np already imported)
# Re-use synthetic regression data
X_reg_et, y_reg_et = make_regression(n_features=10, n_informative=5, random_state=0, shuffle=False)
# Initialize and train Extra Trees Regressor
et_regr = ExtraTreesRegressor(n_estimators=100, max_depth=10, random_state=0)
et_regr.fit(X_reg_et, y_reg_et)
# Make a prediction on the same sample point as RF Regressor example
sample_point_et = X_reg_et[0].reshape(1, -1)
prediction_et = et_regr.predict(sample_point_et)
print(f"ET Regressor parameters: {et_regr.get_params()['n_estimators']} estimators, max_depth {et_regr.get_params()['max_depth']}")
print(f"ET Prediction for sample point {sample_point_et}: {prediction_et}")
print(f"Actual value for sample point: {y_reg_et[0]}")
ET Regressor parameters: 100 estimators, max_depth 10
ET Prediction for sample point [[ 0.18500369 0.89837401 0.75398021 -0.40579343 0.01871914 -0.7394844
-0.84078684 -0.46938239 0.17930832 0.00304159]]: [-44.47586897]
Actual value for sample point: -45.04868412282407

随机决策树功能多样,通常只需最少的超参数调优即可提供出色的性能。有关更多详细信息,请参阅 Scikit-Learn 关于集成方法的文档。