Scikit Learn - 支持向量机
Scikit-Learn - 支持向量机 (SVMs)
Section titled “Scikit-Learn - 支持向量机 (SVMs)”本章深入探讨 Support Vector Machines (SVMs),这是一组强大的监督学习模型,用于分类、回归和离群点检测。
SVMs 简介
Section titled “SVMs 简介”Support Vector Machines 是用途广泛且有效的机器学习算法。它们特别适合对复杂但规模小或中等的数据集进行分类。SVMs 内存效率高,因为它们在决策函数中只使用训练点的子集(即 Support Vectors)。
SVM 用于分类的核心思想是在 N 维空间(其中 N 是特征数量)中找到一个能清晰地将数据点分类的 Hyperplane(超平面)。对于一个线性可分的数据集,SVM 旨在找到最大化 Margin(间隔)的 Hyperplane,即 Hyperplane 与来自任何类别的最近数据点之间的距离。这被称为 Maximum Marginal Hyperplane (MMH),最大间隔超平面。
SVM 中的关键概念:
- Support Vectors(支持向量):这些是离 Hyperplane 最近的数据点。它们对于定义 Hyperplane 的位置和方向至关重要。如果移动这些点,Hyperplane 也会移动。
- Hyperplane(超平面):在 2D 空间中,它是一条直线;在 3D 空间中,它是一个平面;在更高维度中,它是一个超平面。它是分隔类别的决策边界。
- Margin(间隔):Hyperplane 与来自任一类别的最近的支持向量之间的距离。SVMs 试图最大化这个 Margin。
想象一个具有两类点的 2D 散点图。SVM 会试图找到一条直线,用尽可能大的“街道”或“间隙”来分隔这两类点。这条街道的边缘由支持向量定义。
Scikit-Learn 的 SVM 实现支持密集型(NumPy arrays)和稀疏型(SciPy sparse matrices)输入数据。
用于分类的 SVM
Section titled “用于分类的 SVM”Scikit-Learn 提供了三个主要的 SVM 分类类:SVC、NuSVC 和 LinearSVC。
SVC (sklearn.svm.SVC)
Section titled “SVC (sklearn.svm.SVC)”SVC (C-Support Vector Classification) 基于 libsvm 库。它可以执行二分类和多分类。对于多分类问题,它默认使用 one-vs-one(一对一)方案(decision_function_shape='ovo'),不过 decision_function_shape='ovr'(one-vs-rest,一对多)也可用,并且通常为了与其他分类器保持一致而更受欢迎。
SVC 的主要参数
Section titled “SVC 的主要参数”| Parameter | Description |
|---|---|
| C | float, default=1.0。正则化参数。正则化的强度与 C 成反比。必须是严格正值。误差项的惩罚参数 C。 |
| kernel | {‘linear’, ‘poly’, ‘rbf’, ‘sigmoid’, ‘precomputed’},default=‘rbf’。指定算法中使用的核函数类型。‘rbf’(径向基函数)是一个常见选择。 |
| degree | int, default=3。多项式核函数(‘poly’)的次数。被所有其他核函数忽略。 |
| gamma | {‘scale’, ‘auto’} 或 float, default=‘scale’。‘rbf’、‘poly’ 和 ‘sigmoid’ 核函数的系数。如果 gamma='scale'(默认),则使用 1 / (n_features * X.var())。如果 'auto',则使用 1 / n_features。 |
| coef0 | float, default=0.0。核函数中的独立项。仅在 ‘poly’ 和 ‘sigmoid’ 中有意义。 |
| shrinking | bool, default=True。是否使用缩减启发式算法,这可以加快训练速度。 |
| probability | bool, default=False。是否启用概率估计。必须在调用 fit 之前启用,并且会减慢训练速度。 |
| tol | float, default=1e-3。停止标准的容差。 |
| cache_size | float, default=200.0。指定核函数缓存的大小(以 MB 为单位)。 |
| class_weight | dict 或 ‘balanced’,default=None。将类别 i 的参数 C 设置为 class_weight[i]*C。如果为 ‘balanced’,权重与类别频率成反比。 |
| verbose | bool, default=False。启用详细输出。 |
| max_iter | int, default=-1。求解器内部迭代的硬限制,或 -1 表示无限制。 |
| decision_function_shape | {‘ovo’, ‘ovr’},default=‘ovr’。是返回形状为 (n_samples, n_classes) 的一对多 (‘ovr’) 决策函数(与所有其他分类器相同),还是返回形状为 (n_samples, n_classes * (n_classes - 1) / 2) 的原始一对一 (‘ovo’) 决策函数。 |
| break_ties | bool, default=False。如果为 True,predict 将根据 decision_function 的置信度值来打破平局;否则返回平局类别中的第一个类别。 |
| random_state | int, RandomState 实例或 None,default=None。当 probability=True 时,控制用于打乱数据以进行概率估计的伪随机性。 |
SVC 的主要属性
Section titled “SVC 的主要属性”| Attribute | Description |
|---|---|
| support_ | shape 为 (n_SV,) 的 array-like。Support vectors(支持向量)的索引。 |
| support_vectors_ | shape 为 (n_SV, n_features) 的 array-like。Support vectors(支持向量)。 |
| n_support_ | shape 为 (n_classes,),dtype 为 int32 的 array-like。每个类别的 Support vectors(支持向量)数量。 |
| dual_coef_ | array, shape 为 (n_classes-1, n_SV)。决策函数中 Support vectors(支持向量)的系数(用于二分类和多分类 ovo)。 |
| coef_ | array, shape 为 (n_classes * (n_classes-1) / 2, n_features) 或 (1, n_features) 用于二分类。分配给特征的权重(原始问题中的系数)。仅在 kernel='linear' 时可用。 |
| intercept_ | array, shape 为 (n_classes * (n_classes-1) / 2,) 或 (1,) 用于二分类。决策函数中的常数项/截距。 |
| classes_ | shape 为 (n_classes,) 的 array。唯一的类别标签。 |
实现示例 (SVC)
Section titled “实现示例 (SVC)”import numpy as npfrom sklearn.svm import SVCfrom sklearn.preprocessing import StandardScalerfrom sklearn.pipeline import make_pipeline
# Sample data (linearly separable for simplicity with linear kernel)X = np.array([[-1, -1], [-2, -1], [1, 1], [2, 1], [-1, 2], [2, -2]])y = np.array([1, 1, 2, 2, 1, 2]) # Two classes
# It's good practice to scale data for SVMssvc_pipeline = make_pipeline(StandardScaler(), SVC(kernel='linear', C=1.0, random_state=42))svc_pipeline.fit(X, y)
# Making a predictionnew_point = np.array([[-0.8, -0.8]])prediction = svc_pipeline.predict(new_point)print(f"Prediction for {new_point}: {prediction}")
# Accessing attributes from the SVC model within the pipelinesvc_model = svc_pipeline.named_steps['svc']print(f"Support vectors:\n{svc_model.support_vectors_}")if svc_model.kernel == 'linear': print(f"Coefficients (w): {svc_model.coef_}")print(f"Intercept (b): {svc_model.intercept_}")输出示例 (SVC):
Section titled “输出示例 (SVC):”Prediction for [[-0.8 -0.8]]: [1]Support vectors:[[-1. -1.] [-1. 2.] [ 1. 1.]]Coefficients (w): [[ 0.72239822 -0.18059955]]Intercept (b): [0.03611991]NuSVC (sklearn.svm.NuSVC)
Section titled “NuSVC (sklearn.svm.NuSVC)”NuSVC (Nu-Support Vector Classification) 类似于 SVC,但使用参数 nu 来控制支持向量的数量和训练误差。nu 是训练误差比例的上界和支持向量比例的下界。它应该在 (0, 1] 的区间内。
NuSVC 的主要参数 nu
Section titled “NuSVC 的主要参数 nu”nu: float, default=0.5。替代 SVC 的 C 参数用于控制模型复杂度。
实现示例 (NuSVC)
Section titled “实现示例 (NuSVC)”from sklearn.svm import NuSVC# (Using X, y, StandardScaler, make_pipeline from SVC example)
nusvc_pipeline = make_pipeline(StandardScaler(), NuSVC(kernel='rbf', nu=0.5, gamma='scale', random_state=42))nusvc_pipeline.fit(X, y)
prediction_nusvc = nusvc_pipeline.predict(new_point) # Using new_point from SVC exampleprint(f"NuSVC Prediction for {new_point}: {prediction_nusvc}")
nusvc_model = nusvc_pipeline.named_steps['nusvc']print(f"Number of support vectors for each class in NuSVC: {nusvc_model.n_support_}")输出示例 (NuSVC):
Section titled “输出示例 (NuSVC):”NuSVC Prediction for [[-0.8 -0.8]]: [1]Number of support vectors for each class in NuSVC: [3 3]LinearSVC (sklearn.svm.LinearSVC)
Section titled “LinearSVC (sklearn.svm.LinearSVC)”LinearSVC (Linear Support Vector Classification) 是专门用于线性核函数的 SVM 实现。它基于 liblinear 而不是 libsvm(SVC 和 NuSVC 使用的是 libsvm)。LinearSVC 对于大型数据集通常更快,并且在选择惩罚项(L1/L2)和损失函数(‘hinge’, ‘squared_hinge’)方面提供了更大的灵活性。它没有 kernel 参数(因为它总是线性的)。
LinearSVC 的主要参数
Section titled “LinearSVC 的主要参数”与 SVC 的区别:
- penalty: {‘l1’, ‘l2’},default=‘l2’。指定惩罚中使用的范数。
- loss: {‘hinge’, ‘squared_hinge’},default=‘squared_hinge’。指定损失函数。‘hinge’ 是标准的 SVM 损失。
- dual: bool 或 ‘auto’,default=‘auto’。选择算法是求解对偶优化问题还是原始优化问题。当
n_samples > n_features时优先使用dual=False。‘auto’ 会自动进行选择。 - 没有
kernel、gamma、degree、coef0、shrinking、probability参数。 - 缺少
support_、support_vectors_、n_support_、dual_coef_属性。
实现示例 (LinearSVC)
Section titled “实现示例 (LinearSVC)”from sklearn.svm import LinearSVCfrom sklearn.datasets import make_classification# (StandardScaler, make_pipeline already imported)
# Generate a larger dataset for LinearSVC demonstrationX_lin, y_lin = make_classification(n_samples=200, n_features=4, random_state=0)
linearsvc_pipeline = make_pipeline(StandardScaler(), LinearSVC(penalty='l2', loss='squared_hinge', dual='auto', random_state=0, tol=1e-4, C=1.0))linearsvc_pipeline.fit(X_lin, y_lin)
# Making a prediction on a sample pointsample_lin_point = X_lin[0].reshape(1, -1)prediction_linearsvc = linearsvc_pipeline.predict(sample_lin_point)print(f"LinearSVC Prediction for {sample_lin_point[0,:2]}...: {prediction_linearsvc}")
linearsvc_model = linearsvc_pipeline.named_steps['linearsvc']print(f"LinearSVC Coefficients: {linearsvc_model.coef_}")print(f"LinearSVC Intercept: {linearsvc_model.intercept_}")输出示例 (LinearSVC):
Section titled “输出示例 (LinearSVC):”LinearSVC Prediction for [0.16762098 0.20222924]...: [1]LinearSVC Coefficients: [[ 0.15689569 -0.24839626 0.75454966 0.41789389]]LinearSVC Intercept: [0.16693748]用于回归的 SVM (SVR)
Section titled “用于回归的 SVM (SVR)”SVM 的概念可以扩展到解决回归问题。这被称为 Support Vector Regression (SVR),支持向量回归。在 SVR 中,目标是找到一个函数(Hyperplane),使得尽可能多的训练点与 y 的偏差不超过一个 Margin epsilon,同时该函数尽可能地平坦。在预测值周围 epsilon 范围内的点不会计入损失。
Scikit-Learn 提供了 SVR、NuSVR 和 LinearSVR。
SVR (sklearn.svm.SVR)
Section titled “SVR (sklearn.svm.SVR)”Epsilon-Support Vector Regression,基于 libsvm。主要参数是 C(正则化)和 epsilon(定义不产生误差惩罚的容差范围)。
SVR 的主要参数 epsilon
Section titled “SVR 的主要参数 epsilon”epsilon: float, default=0.1。指定了在训练损失函数中,预测值与实际值之间距离在 epsilon 范围内不会产生惩罚的 epsilon 范围。
实现示例 (SVR)
Section titled “实现示例 (SVR)”from sklearn.svm import SVRfrom sklearn.datasets import make_regression# (StandardScaler, make_pipeline, np already imported)
X_reg, y_reg = make_regression(n_samples=100, n_features=1, random_state=42, noise=5)
svr_pipeline = make_pipeline(StandardScaler(), SVR(kernel='rbf', C=100, gamma=0.1, epsilon=0.1))svr_pipeline.fit(X_reg, y_reg)
sample_reg_point = np.array([[1.5]]) # Example new data pointprediction_svr = svr_pipeline.predict(sample_reg_point)print(f"SVR Prediction for {sample_reg_point}: {prediction_svr}")# SVR attributes like support_vectors_ are also available输出示例 (SVR):
Section titled “输出示例 (SVR):”SVR Prediction for [[1.5]]: [66.9114178]NuSVR (sklearn.svm.NuSVR)
Section titled “NuSVR (sklearn.svm.NuSVR)”Nu-Support Vector Regression。类似于 SVR,但使用 nu 来控制支持向量的数量,而不是使用 epsilon 来控制不敏感区域。
实现示例 (NuSVR)
Section titled “实现示例 (NuSVR)”from sklearn.svm import NuSVR# (Using X_reg, y_reg, StandardScaler, make_pipeline, np from SVR example)
nusvr_pipeline = make_pipeline(StandardScaler(), NuSVR(kernel='rbf', C=100, nu=0.1, gamma='scale'))nusvr_pipeline.fit(X_reg, y_reg)
prediction_nusvr = nusvr_pipeline.predict(sample_reg_point)print(f"NuSVR Prediction for {sample_reg_point}: {prediction_nusvr}")输出示例 (NuSVR):
Section titled “输出示例 (NuSVR):”NuSVR Prediction for [[1.5]]: [66.6761768]LinearSVR (sklearn.svm.LinearSVR)
Section titled “LinearSVR (sklearn.svm.LinearSVR)”Linear Support Vector Regression,基于 liblinear。对于大型数据集上的线性 SVR 更快。主要参数是 epsilon。loss 可以是 'epsilon_insensitive' 或 'squared_epsilon_insensitive'。
实现示例 (LinearSVR)
Section titled “实现示例 (LinearSVR)”from sklearn.svm import LinearSVR# (Using X_reg, y_reg, StandardScaler, make_pipeline, np from SVR example)
linearsvr_pipeline = make_pipeline(StandardScaler(), LinearSVR(epsilon=0.0, C=1.0, loss='squared_epsilon_insensitive', dual='auto', random_state=0, tol=1e-4))linearsvr_pipeline.fit(X_reg, y_reg)
prediction_linearsvr = linearsvr_pipeline.predict(sample_reg_point)print(f"LinearSVR Prediction for {sample_reg_point}: {prediction_linearsvr}")
linearsvr_model = linearsvr_pipeline.named_steps['linearsvr']print(f"LinearSVR Coefficients: {linearsvr_model.coef_}")print(f"LinearSVR Intercept: {linearsvr_model.intercept_}")输出示例 (LinearSVR):
Section titled “输出示例 (LinearSVR):”LinearSVR Prediction for [[1.5]]: [67.53298058]LinearSVR Coefficients: [45.66113378]LinearSVR Intercept: [-0.95663479]SVMs 是强大的模型。选择正确的 Kernel(核函数)和参数(如 C、gamma、nu、epsilon)至关重要且通常需要实验(例如,使用 GridSearchCV)。强烈建议进行特征缩放。更多信息,请访问 Scikit-Learn SVM 文档。