Skip to content

监督学习:分类

使用 Python 进行 AI 开发 – 监督学习:分类

Section titled “使用 Python 进行 AI 开发 – 监督学习:分类”

本章重点介绍监督学习(Supervised Learning),特别是分类(Classification)任务。分类是一种监督机器学习(Machine Learning)类型,其目标是为给定的输入预测一个分类类别标签(Categorical Class Label)。

分类模型从观测数据(特征,Features)及其对应的真实标签(Labels)中学习,以便对新的、未见过的数据进行预测。输出是一个离散的类别,例如“垃圾邮件”或“非垃圾邮件”,“猫”或“狗”,或医疗诊断结果如“恶性”或“良性”。构建分类模型需要一个带有标签的训练数据集(Labeled Training Dataset)。常见的应用包括图像识别、医疗诊断、情感分析和欺诈检测。

使用 Scikit-learn 在 Python 中构建分类器的步骤

Section titled “使用 Scikit-learn 在 Python 中构建分类器的步骤”

我们将使用 Python 3 和 Scikit-learn 库,这是一个强大且流行的 Python 机器学习工具包。以下是一般的工作流程:

如果您尚未安装 Scikit-learn,请进行安装。最常用的方法是使用 pip:

pip install scikit-learn

如果您使用 Anaconda,可以通过 conda 安装:

conda install scikit-learn

Scikit-learn 自带了一些内置数据集。对于本例,我们将使用“威斯康星乳腺癌诊断数据库”(Breast Cancer Wisconsin Diagnostic Database)。该数据集包含乳腺癌肿瘤的信息,带有将其分类为恶性(malignant)或良性(benign)的标签。

该数据集有 569 个实例(肿瘤)和 30 个特征(属性,如半径、纹理、平滑度、面积等)。

from sklearn.datasets import load_breast_cancer
# Load the dataset
cancer_data = load_breast_cancer()

Scikit-learn 数据集通常是 Bunch 对象,它们类似于字典。关键属性包括:

  • data:特征矩阵(NumPy 数组)。
  • target:标签数组(NumPy 数组)。
  • feature_names:特征的名称。
  • target_names:目标类别的名称。

为了清晰起见,我们将这些赋值给变量:

features = cancer_data.data
labels = cancer_data.target
feature_names = cancer_data.feature_names
target_names = cancer_data.target_names

我们来查看一些数据:

print("Target names:", target_names)
print("First 5 labels:", labels[:5]) # 0 for malignant, 1 for benign
print("First feature name:", feature_names[0])
print("First instance's features:\n", features[0])

示例输出:

Target names: ['malignant' 'benign']
First 5 labels: [0 0 0 0 0]
First feature name: mean radius
First instance's features:
[1.799e+01 1.038e+01 1.228e+02 1.001e+03 1.184e-01 2.776e-01
3.001e-01 1.471e-01 2.419e-01 7.871e-02 1.095e+00 9.053e-01
8.589e+00 1.534e+02 6.399e-03 4.904e-02 5.373e-02 1.587e-02
3.003e-02 6.193e-03 2.538e+01 1.733e+01 1.846e+02 2.019e+03
1.622e-01 6.656e-01 7.119e-01 2.654e-01 4.601e-01 1.189e-01]

这表明第一个肿瘤被标记为“恶性”(0),其“平均半径”约为 17.99。

步骤 3:将数据分割为训练集和测试集

Section titled “步骤 3:将数据分割为训练集和测试集”

在模型训练期间,从未见过的数据上评估模型至关重要。我们将数据集分割为训练集(用于模型学习)和测试集(用于模型评估)。Scikit-learn 的 train_test_split 函数用于此目的。

from sklearn.model_selection import train_test_split
# Split data: 70% for training, 30% for testing.
# random_state ensures reproducibility.
X_train, X_test, y_train, y_test = train_test_split(
features, labels, test_size=0.30, random_state=42
)
print(f"Training set size: {X_train.shape[0]} samples")
print(f"Test set size: {X_test.shape[0]} samples")

常见的分割比例是 70-80% 用于训练,20-30% 用于测试。

我们将使用高斯朴素贝叶斯(Gaussian Naïve Bayes,GNB)算法作为我们的第一个分类器。它是一种简单但通常有效的算法,基于贝叶斯定理(Bayes’ Theorem),并假设特征之间相互独立。

from sklearn.naive_bayes import GaussianNB
# Initialize the Gaussian Naïve Bayes classifier
gnb_model = GaussianNB()
# Train the model using the training data
gnb_model.fit(X_train, y_train)
print("Model training complete.")

训练后,我们使用测试集评估模型的性能。predict() 方法对测试集的特征进行预测。

# Make predictions on the test set
y_pred = gnb_model.predict(X_test)
print("First 10 predictions:", y_pred[:10])
print("First 10 actual labels:", y_test[:10])

为了量化性能,我们使用评估指标。准确率(Accuracy)是一个常见的指标:

from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print(f"Model Accuracy: {accuracy * 100:.2f}%")

对于该数据集和 GNB,准确率通常在 90-95% 左右。例如,输出可能是 Model Accuracy: 94.15%。

这种结构化方法——加载数据、预处理(如果需要)、分割、训练、评估——是构建机器学习模型的基础。

让我们探讨一些 Scikit-learn 中更广泛使用的分类算法。

朴素贝叶斯分类器(Naïve Bayes Classifiers)是一系列简单的概率分类器,它们基于贝叶斯定理,并对特征之间进行强(朴素)独立性假设。尽管它们很简单,但性能可能很好,尤其是在高维数据(如文本)上。

Scikit-learn 提供了几种类型:GaussianNB(用于假设高斯分布的连续特征,如上所示)、MultinomialNB(用于离散计数,例如文本中的词频)和 BernoulliNB(用于二元/布尔特征)。

前面的示例展示了在乳腺癌数据集上使用 GaussianNB。步骤保持不变:初始化模型,使用训练数据(X_train, y_train)进行拟合,对测试数据(X_test)进行预测,然后评估。

支持向量机(Support Vector Machines,SVMs)是强大且通用的监督学习模型,用于分类、回归和异常检测。对于分类任务,SVM 的目标是在 N 维空间(N 是特征数量)中找到一个最佳超平面(Hyperplane),以分隔不同类别的数据点。

“支持向量”(Support Vectors)是距离超平面最近的数据点,如果移除它们,将改变超平面的位置。SVM 可以使用各种核函数(Kernels)(例如,线性核、多项式核、RBF 核、Sigmoid 核)处理线性可分和非线性可分数据。

我们来使用 Iris 数据集构建一个 SVM 分类器。该数据集包含 3 类 Iris 植物和 4 个特征(萼片长度/宽度,花瓣长度/宽度)。为了便于可视化,我们将只使用前两个特征。

import numpy as np
import matplotlib.pyplot as plt
from sklearn import svm, datasets
# Load Iris dataset
iris = datasets.load_iris()
X = iris.data[:, :2] # We only take the first two features for 2D visualization
y = iris.target
# Create an SVM Classifier with a linear kernel
# C is the regularization parameter
C = 1.0
svc_model = svm.SVC(kernel='linear', C=C, decision_function_shape='ovr').fit(X, y)
# Create a mesh to plot the decision boundaries
x_min, x_max = X[:, 0].min() - 0.5, X[:, 0].max() + 0.5
y_min, y_max = X[:, 1].min() - 0.5, X[:, 1].max() + 0.5
h = 0.02 # step size in the mesh
xx, yy = np.meshgrid(np.arange(x_min, x_max, h),
np.arange(y_min, y_max, h))
# Plot decision boundary. For that, we will assign a color to each point in the mesh [x_min, x_max]x[y_min, y_max].
Z = svc_model.predict(np.c_[xx.ravel(), yy.ravel()])
Z = Z.reshape(xx.shape)
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, Z, cmap=plt.cm.coolwarm, alpha=0.8)
# Plot also the training points
plt.scatter(X[:, 0], X[:, 1], c=y, cmap=plt.cm.coolwarm, edgecolors='k')
plt.xlabel('Sepal length')
plt.ylabel('Sepal width')
plt.xlim(xx.min(), xx.max())
plt.ylim(yy.min(), yy.max())
plt.xticks(())
plt.yticks(())
plt.title('SVC with linear kernel on Iris dataset (first two features)')
plt.show()

这段代码将显示 Iris 数据(萼片长度 vs 萼片宽度)的散点图,不同类别用不同颜色表示。由线条(决策边界,Decision Boundaries)分隔的区域表示 SVM 模型预测的分类区域。

像 ‘rbf’ (径向基函数) 这样的核函数可用于非线性可分数据:svm.SVC(kernel='rbf', C=C, gamma='auto')。

尽管名字里带有“回归”,但逻辑回归(Logistic Regression)是一种广泛用于二元分类(Binary Classification)和多类别分类(Multi-class Classification)的算法。它建模了给定输入点属于特定类别的概率。然后使用逻辑函数(或 Sigmoid 函数)将这个概率转换为 0 到 1 之间的值。

因变量是我们预测的目标类别,自变量是特征。例如,基于年龄、浏览历史等特征,预测客户是否会点击广告(二元:是/否)。

逻辑回归示例:

import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
# Generate synthetic data for classification
X_synthetic, y_synthetic = make_classification(
n_samples=100, n_features=2, n_redundant=0,
n_informative=2, random_state=1, n_clusters_per_class=1
)
# Create and train the Logistic Regression classifier
# 'solver' and 'C' are important hyperparameters
# 'liblinear' is good for small datasets, 'lbfgs' is a common default
log_reg_model = LogisticRegression(solver='liblinear', C=1.0, random_state=42)
log_reg_model.fit(X_synthetic, y_synthetic)
# Function to visualize decision boundaries (similar to SVM example)
def plot_decision_boundary(model, X, y, title):
x_min, x_max = X[:, 0].min() - 0.5, X[:, 0].max() + 0.5
y_min, y_max = X[:, 1].min() - 0.5, X[:, 1].max() + 0.5
h = .02 # step size in the mesh
xx, yy = np.meshgrid(np.arange(x_min, x_max, h),
np.arange(y_min, y_max, h))
Z = model.predict(np.c_[xx.ravel(), yy.ravel()])
Z = Z.reshape(xx.shape)
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, Z, cmap=plt.cm.RdYlBu, alpha=0.8)
plt.scatter(X[:, 0], X[:, 1], c=y, cmap=plt.cm.RdYlBu, edgecolors='k')
plt.title(title)
plt.xlabel('Feature 1')
plt.ylabel('Feature 2')
plt.show()
plot_decision_boundary(log_reg_model, X_synthetic, y_synthetic, 'Logistic Regression Decision Boundary')

该图将显示数据点和逻辑回归模型学习到的决策边界,在合成数据集中分隔两个类别。

决策树(Decision Tree)是一种用于分类和回归的非参数监督学习方法。它将决策建模为树状结构。每个内部节点(Internal Node)代表对一个属性(Attribute)(例如,颜色是红色吗?)的“测试”,每个分支(Branch)代表测试的结果,每个叶节点(Leaf Node)代表一个类别标签(在计算所有属性后做出的决策)。

示例:基于身高和头发长度预测性别(一个模拟数据集)。

为了可视化树结构本身(而不仅仅是边界),可以使用 graphviz。确保已安装它(pip install graphviz pydotplus)。您还需要安装 Graphviz 软件 (graphviz.org)。

from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
import numpy as np
# Toy dataset: [height (cm), hair length (cm)]
X_gender = np.array([
[165, 19], [175, 32], [136, 35], [174, 65], [141, 28],
[176, 15], [131, 32], [166, 6], [128, 32], [179, 10],
[136, 34], [186, 2], [126, 25], [176, 28], [112, 38],
[169, 9], [171, 36], [116, 25], [196, 25]
])
# Labels: 0 for 'Man', 1 for 'Woman'
Y_gender = np.array([0, 1, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 1, 1, 0, 1, 1, 0])
data_feature_names = ['height', 'length of hair']
class_names_gender = ['Man', 'Woman']
# Split data (optional for this small dataset, but good practice)
X_train_g, X_test_g, y_train_g, y_test_g = train_test_split(
X_gender, Y_gender, test_size=0.30, random_state=5
)
# Create and fit the Decision Tree model
dt_model = DecisionTreeClassifier(random_state=42, max_depth=3) # max_depth limits tree size
dt_model.fit(X_train_g, y_train_g)
# Make a prediction
sample_pred = dt_model.predict([[133, 37]])
print(f"Prediction for [133cm, 37cm hair]: {class_names_gender[sample_pred[0]]}")
# Display the tree structure as text
tree_rules = export_text(dt_model, feature_names=data_feature_names, class_names=class_names_gender)
print("\nDecision Tree Rules:\n", tree_rules)
# Evaluate accuracy
accuracy_dt = dt_model.score(X_test_g, y_test_g)
print(f"Decision Tree Accuracy: {accuracy_dt*100:.2f}%")

输出将包含一个预测结果(例如,“Woman”)以及决策树规则的文本表示,显示如何使用特征进行分割。例如:

Prediction for [133cm, 37cm hair]: Woman
Decision Tree Rules:
|--- length of hair <= 26.50
| |--- height <= 170.00
| | |--- class: Man
| |--- height > 170.00
| | |--- class: Man
|--- length of hair > 26.50
| |--- height <= 140.00
| | |--- class: Woman
| |--- height > 140.00
| | |--- class: Woman
Decision Tree Accuracy: 83.33%

如果您安装了 graphviz 和 pydotplus,可以将其可视化为图像:from sklearn.tree import export_graphviz; import pydotplus; dot_data = export_graphviz(...); graph = pydotplus.graph_from_dot_data(dot_data); graph.write_png('decision_tree.png')。

随机森林(Random Forest)是一种集成学习(Ensemble Learning)方法,它通过在训练时构建多个决策树来工作。对于分类任务,输出是大多数树选择的类别。它通常比单个决策树更鲁棒和准确,通过平均结果减少过拟合(Overfitting)。

我们将再次使用乳腺癌数据集。

from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.datasets import load_breast_cancer
import matplotlib.pyplot as plt
import numpy as np
# Load data
cancer = load_breast_cancer()
X_rf, y_rf = cancer.data, cancer.target
# Split data
X_train_rf, X_test_rf, y_train_rf, y_test_rf = train_test_split(X_rf, y_rf, random_state=0)
# Create and train Random Forest model
# n_estimators is the number of trees in the forest
rf_model = RandomForestClassifier(n_estimators=100, random_state=0, max_depth=5)
rf_model.fit(X_train_rf, y_train_rf)
# Evaluate accuracy
acc_train_rf = rf_model.score(X_train_rf, y_train_rf)
acc_test_rf = rf_model.score(X_test_rf, y_test_rf)
print(f'Random Forest - Training Accuracy: {acc_train_rf*100:.2f}%')
print(f'Random Forest - Test Accuracy: {acc_test_rf*100:.2f}%')
# Feature importances
importances = rf_model.feature_importances_
indices = np.argsort(importances)[::-1] # Sort features by importance
print("\nTop 10 Feature Importances:")
for i in range(10):
print(f"{i+1}. Feature '{cancer.feature_names[indices[i]]}': {importances[indices[i]]:.4f}")
# Plotting feature importances (textual description)
plt.figure(figsize=(10,8))
plt.title('Feature Importances in Random Forest')
plt.barh(range(10), importances[indices][:10][::-1], align='center') # Plot top 10, reversed for plot order
plt.yticks(range(10), cancer.feature_names[indices][:10][::-1])
plt.xlabel('Relative Importance')
plt.tight_layout()
plt.show()

输出将显示训练集和测试集上的准确率(例如,训练集:约 99-100%,测试集:约 96-97%)。它还将列出模型识别出的最重要的特征。图表将把这些特征重要性可视化为水平条形图,较长的条形表示更重要的特征,例如 ‘worst concave points’、‘worst perimeter’、‘worst radius’。

训练分类器后,我们需要评估其有效性。会使用几种指标,这些指标通常源自混淆矩阵(Confusion Matrix)。

混淆矩阵是一个用于描述分类模型在已知真实值的一组测试数据上性能的表格。它比较机器学习模型预测的值与实际目标值。

对于二元分类问题(例如,正类/负类 或 1/0):

预测值: 正类 (1) 负类 (0)

实际值 正类 (1): TP (真正例) FN (假反例)

实际值 负类 (0): FP (假正例) TN (真反例)

  • 真正例 (TP):正确预测为正类的实例。
  • 真反例 (TN):正确预测为负类的实例。
  • 假正例 (FP):错误预测为正类的实例(第一类错误,Type I error)。
  • 假反例 (FN):错误预测为负类的实例(第二类错误,Type II error)。
from sklearn.metrics import confusion_matrix, classification_report
# Using the GaussianNB model from the first example
# y_test, y_pred were calculated earlier
cm = confusion_matrix(y_test, y_pred)
print("Confusion Matrix (GNB Model):")
print(cm)
# For more detailed metrics:
print("\nClassification Report (GNB Model):")
print(classification_report(y_test, y_pred, target_names=target_names))

GNB 模型在乳腺癌数据上的混淆矩阵示例输出:

Confusion Matrix (GNB Model):
[[ 59 5]
[ 5 102]]
This means: TP=102 (benign correctly id'd), TN=59 (malignant correctly id'd), FP=5 (malignant misclassified as benign), FN=5 (benign misclassified as malignant). (Assuming benign is positive class)

分类报告(Classification Report)提供了每个类别的精确率(Precision)、召回率(Recall)、F1 分数(F1-Score)和支持度(Support)。

在所有被评估的案例中,正确预测所占的比例。

$$Accuracy = \frac{TP+TN}{TP+FP+FN+TN}$$

在所有预测为正类的实例中,实际为正类的比例是多少?高精确率意味着低假正例率。

$$Precision = \frac{TP}{TP+FP}$$

召回率 (Recall)(敏感度 Sensitivity 或 真正例率 True Positive Rate)

Section titled “召回率 (Recall)(敏感度 Sensitivity 或 真正例率 True Positive Rate)”

在所有实际为正类的实例中,模型正确识别的比例是多少?高召回率意味着低假反例率。

$$Recall = \frac{TP}{TP+FN}$$

特异度 (Specificity)(真反例率 True Negative Rate)

Section titled “特异度 (Specificity)(真反例率 True Negative Rate)”

在所有实际为负类的实例中,模型正确识别的比例是多少?

$$Specificity = \frac{TN}{TN+FP}$$

精确率和召回率的调和平均数。当您需要在精确率和召回率之间取得平衡时,F1 分数特别有用,尤其是在类别分布不均的情况下。

$$F1\ Score = 2 * \frac{Precision * Recall}{Precision + Recall}$$

类别不平衡(Class Imbalance)发生在一个类别的观测数量(少数类,Minority Class)显著低于或高于(多数类,Majority Class)其他类别时。这在现实世界的问题中很常见,例如欺诈检测、罕见疾病识别或制造业中的缺陷检测。

在不平衡数据上训练的标准分类器通常在少数类上表现不佳,因为它们偏向于多数类(这样做可以最大化总体准确率)。

数据集:信用卡欺诈检测
总交易次数 = 100,000
欺诈交易(少数类) = 100 (0.1%)
非欺诈交易(多数类) = 99,900 (99.9%)
事件发生率(欺诈) = 0.1%

一个总是预测“非欺诈”的模型可以达到 99.9% 的准确率,但对于检测欺诈来说是无用的。

1. 重采样技术 (Resampling Techniques)

Section titled “1. 重采样技术 (Resampling Techniques)”

重采样的目标是平衡训练数据集中的类别分布。

  • 随机欠采样(Random Under-Sampling):随机移除多数类的实例。这会减小数据集大小,可能会丢失有用的信息。
  • 随机过采样(Random Over-Sampling):随机复制少数类的实例。这会增加数据集大小,可能导致对少数类过拟合。
  • 合成少数类过采样技术(Synthetic Minority Over-sampling Technique,SMOTE):通过在现有少数类实例之间进行插值来创建合成样本。比随机过采样更复杂。
  • ADASYN(自适应合成采样,Adaptive Synthetic Sampling):类似于 SMOTE,但为更难学习的少数类样本生成更多合成数据。

imbalanced-learn 库(pip install imbalanced-learn)提供了这些技术的实现(例如,RandomUnderSampler、RandomOverSampler、SMOTE)。

2. 算法方法 (Algorithmic Approaches)(成本敏感学习 Cost-Sensitive Learning)

Section titled “2. 算法方法 (Algorithmic Approaches)(成本敏感学习 Cost-Sensitive Learning)”

修改现有算法以考虑类别不平衡。这通常涉及为少数类分配更高的错误分类成本。许多 Scikit-learn 分类器(例如,SVC、LogisticRegression、RandomForestClassifier)都有一个 class_weight 参数,可以设置为 'balanced' 或自定义的权重字典。

# Example with RandomForestClassifier
rf_balanced_model = RandomForestClassifier(n_estimators=100, class_weight='balanced', random_state=0)
# Or custom weights: class_weight={0:1, 1:10} if class 1 is minority
# Then train and evaluate as usual.

随机森林或梯度提升(Gradient Boosting)等集成方法可以针对不平衡数据进行调整。平衡随机森林(Balanced Random Forest)或 EasyEnsemble/BalanceCascade 等技术专门解决不平衡问题。

准确率对于不平衡数据集具有误导性。应关注诸如精确率、召回率、F1 分数(特别是针对少数类)、ROC 曲线下面积(Area Under the ROC Curve,AUC-ROC)和精确率-召回率曲线下面积(Area Under the Precision-Recall Curve,AUC-PR)等指标。

解决类别不平衡对于构建有效的实际分类器通常至关重要。