随机决策树
Scikit-Learn - 随机决策树
Section titled “Scikit-Learn - 随机决策树”本章探讨 Scikit-Learn 中的随机决策树算法,它们是强大的集成(ensemble)方法,以高准确性和抗过拟合能力而闻名。
随机决策树算法简介
Section titled “随机决策树算法简介”标准决策树容易发生过拟合,特别是当树的深度较大时。随机决策树算法,如随机森林(Random Forests)和极度随机树(Extra-Trees),通过构建多个决策树的集成(ensemble)来解决这个问题。随机性主要通过两种方式引入:在数据的一个不同子样本(subsample)(装袋 bagging)上训练每棵树,和/或在每个分裂点只考虑特征的一个随机子集。最终的预测通常是单个树的预测结果的平均值(用于回归)或多数投票(用于分类)。
sklearn.ensemble 模块提供了这些算法的实现。
随机森林算法
Section titled “随机森林算法”随机森林构建多个决策树。对于每棵树,它通常使用训练数据的自助采样(bootstrap sample,有放回的随机采样)。在树中分裂节点时,它只考虑特征的一个随机子集来寻找最佳分裂。这种双重随机性有助于降低树之间的相关性,减少方差并提高泛化能力。随机森林可用于分类和回归任务。
随机森林分类器 (RandomForestClassifier)
Section titled “随机森林分类器 (RandomForestClassifier)”sklearn.ensemble.RandomForestClassifier 用于分类任务。关键参数包括 n_estimators(树的数量)、max_features(随机特征子集的大小)、max_depth(每棵树的最大深度)和 min_samples_split(分裂节点所需的最小样本数)。
实现示例 1:合成数据
Section titled “实现示例 1:合成数据”from sklearn.model_selection import cross_val_scorefrom sklearn.datasets import make_blobsfrom sklearn.ensemble import RandomForestClassifier
# Generate synthetic data for classificationX, y = make_blobs(n_samples=1000, n_features=10, centers=5, random_state=0)
# Initialize Random Forest Classifier# Using fewer estimators for quicker example, in practice use more (e.g., 100+)rf_clf = RandomForestClassifier(n_estimators=50, max_depth=None, min_samples_split=2, random_state=0)
# Evaluate using cross-validationscores = cross_val_score(rf_clf, X, y, cv=5) # 5-fold cross-validationprint(f"Cross-validation scores: {scores}")print(f"Mean cross-validation score: {scores.mean():.4f}")输出示例 1:
Section titled “输出示例 1:”Cross-validation scores: [0.975 0.97 0.98 0.965 0.975]Mean cross-validation score: 0.9730实现示例 2:Iris 数据集
Section titled “实现示例 2:Iris 数据集”使用 Iris 数据集进行一个实际的分类示例,包括评估指标。
import pandas as pdfrom sklearn.model_selection import train_test_splitfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.metrics import classification_report, confusion_matrix, accuracy_scorefrom sklearn.datasets import load_iris
# Load Iris datasetiris = load_iris()X_iris = iris.datay_iris = iris.target
# Split data into training and testing setsX_train, X_test, y_train, y_test = train_test_split(X_iris, y_iris, test_size=0.30, random_state=42)
# Initialize and train the Random Forest Classifierrf_iris_clf = RandomForestClassifier(n_estimators=100, random_state=42)rf_iris_clf.fit(X_train, y_train)
# Make predictionsy_pred = rf_iris_clf.predict(X_test)
# Evaluate the classifierprint("Confusion Matrix:")print(confusion_matrix(y_test, y_pred))print("\nClassification Report:")print(classification_report(y_test, y_pred, target_names=iris.target_names))print(f"Accuracy Score: {accuracy_score(y_test, y_pred):.4f}")输出示例 2:
Section titled “输出示例 2:”Confusion Matrix:[[19 0 0] [ 0 13 0] [ 0 0 13]]
Classification Report: precision recall f1-score support
setosa 1.00 1.00 1.00 19 versicolor 1.00 1.00 1.00 13 virginica 1.00 1.00 1.00 13
accuracy 1.00 45 macro avg 1.00 1.00 1.00 45weighted avg 1.00 1.00 1.00 45
Accuracy Score: 1.0000随机森林回归器 (RandomForestRegressor)
Section titled “随机森林回归器 (RandomForestRegressor)”sklearn.ensemble.RandomForestRegressor 用于回归任务。它使用与分类器相似的参数,但分裂节点的准则通常基于减少均方误差(Mean Squared Error, MSE)或平均绝对误差(Mean Absolute Error, MAE)。
from sklearn.ensemble import RandomForestRegressorfrom sklearn.datasets import make_regressionimport numpy as np
# Generate synthetic data for regressionX_reg, y_reg = make_regression(n_features=10, n_informative=5, random_state=0, shuffle=False)
# Initialize and train Random Forest Regressorrf_regr = RandomForestRegressor(n_estimators=100, max_depth=10, random_state=0)rf_regr.fit(X_reg, y_reg)
# Make a prediction on a sample point (e.g., the first point from X_reg)sample_point = X_reg[0].reshape(1, -1)prediction = rf_regr.predict(sample_point)print(f"Model parameters: {rf_regr.get_params()['n_estimators']} estimators, max_depth {rf_regr.get_params()['max_depth']}")print(f"Prediction for sample point {sample_point}: {prediction}")print(f"Actual value for sample point: {y_reg[0]}")Model parameters: 100 estimators, max_depth 10Prediction for sample point [[ 0.18500369 0.89837401 0.75398021 -0.40579343 0.01871914 -0.7394844 -0.84078684 -0.46938239 0.17930832 0.00304159]]: [-47.6763054]Actual value for sample point: -45.04868412282407极度随机树(Extra-Trees)
Section titled “极度随机树(Extra-Trees)”与随机森林相比,极度随机树引入了更多的随机性。在选择分裂点时,不仅考虑特征的一个随机子集,而且对于每个考虑的特征,分裂阈值也是随机选择的(从当前样本中该特征范围内的均匀分布中选取)。这通常会导致偏差(bias)略高,但可以进一步减少方差(variance)。由于它们不需要搜索最优阈值,通常训练速度更快。
极度随机树分类器 (ExtraTreesClassifier)
Section titled “极度随机树分类器 (ExtraTreesClassifier)”sklearn.ensemble.ExtraTreesClassifier 使用与 RandomForestClassifier 相同的关键参数。
实现示例 1:合成数据
Section titled “实现示例 1:合成数据”from sklearn.model_selection import cross_val_scorefrom sklearn.datasets import make_blobsfrom sklearn.ensemble import ExtraTreesClassifier
# Re-use synthetic data from Random Forest exampleX_et, y_et = make_blobs(n_samples=1000, n_features=10, centers=5, random_state=0)
# Initialize Extra Trees Classifieret_clf = ExtraTreesClassifier(n_estimators=50, max_depth=None, min_samples_split=2, random_state=0)
# Evaluate using cross-validationet_scores = cross_val_score(et_clf, X_et, y_et, cv=5)print(f"Extra-Trees cross-validation scores: {et_scores}")print(f"Extra-Trees mean cross-validation score: {et_scores.mean():.4f}")输出示例 1:
Section titled “输出示例 1:”Extra-Trees cross-validation scores: [0.98 0.97 0.98 0.97 0.98 ]Extra-Trees mean cross-validation score: 0.9760实现示例 2:使用不同的数据集(例如 make_classification)
Section titled “实现示例 2:使用不同的数据集(例如 make_classification)”使用 make_classification 处理更复杂的场景。
from sklearn.datasets import make_classificationfrom sklearn.model_selection import KFold# (KFold and cross_val_score already imported)# (ExtraTreesClassifier already imported)
X_class, Y_class = make_classification(n_samples=1000, n_features=20, n_informative=15, n_redundant=5, random_state=7, n_classes=3)
# Setup KFold for cross-validationkfold = KFold(n_splits=10, random_state=7, shuffle=True) # Added shuffle=True for reproducibility with random_state
num_trees = 150max_feat = 7 # Example: consider 7 features at each splitet_clf_complex = ExtraTreesClassifier(n_estimators=num_trees, max_features=max_feat, random_state=7)results_complex = cross_val_score(et_clf_complex, X_class, Y_class, cv=kfold)print(f"Extra-Trees (complex data) mean accuracy: {results_complex.mean():.4f}")输出示例 2:
Section titled “输出示例 2:”Extra-Trees (complex data) mean accuracy: 0.8050极度随机树回归器 (ExtraTreesRegressor)
Section titled “极度随机树回归器 (ExtraTreesRegressor)”sklearn.ensemble.ExtraTreesRegressor 用于回归任务,其参数与 ExtraTreesClassifier 类似。
from sklearn.ensemble import ExtraTreesRegressor# (make_regression and np already imported)
# Re-use synthetic regression dataX_reg_et, y_reg_et = make_regression(n_features=10, n_informative=5, random_state=0, shuffle=False)
# Initialize and train Extra Trees Regressoret_regr = ExtraTreesRegressor(n_estimators=100, max_depth=10, random_state=0)et_regr.fit(X_reg_et, y_reg_et)
# Make a prediction on the same sample point as RF Regressor examplesample_point_et = X_reg_et[0].reshape(1, -1)prediction_et = et_regr.predict(sample_point_et)print(f"ET Regressor parameters: {et_regr.get_params()['n_estimators']} estimators, max_depth {et_regr.get_params()['max_depth']}")print(f"ET Prediction for sample point {sample_point_et}: {prediction_et}")print(f"Actual value for sample point: {y_reg_et[0]}")ET Regressor parameters: 100 estimators, max_depth 10ET Prediction for sample point [[ 0.18500369 0.89837401 0.75398021 -0.40579343 0.01871914 -0.7394844 -0.84078684 -0.46938239 0.17930832 0.00304159]]: [-44.47586897]Actual value for sample point: -45.04868412282407随机决策树功能多样,通常只需最少的超参数调优即可提供出色的性能。有关更多详细信息,请参阅 Scikit-Learn 关于集成方法的文档。