Skip to content

Scikit Learn - 支持向量机

本章深入探讨 Support Vector Machines (SVMs),这是一组强大的监督学习模型,用于分类、回归和离群点检测。

Support Vector Machines 是用途广泛且有效的机器学习算法。它们特别适合对复杂但规模小或中等的数据集进行分类。SVMs 内存效率高,因为它们在决策函数中只使用训练点的子集(即 Support Vectors)。

SVM 用于分类的核心思想是在 N 维空间(其中 N 是特征数量)中找到一个能清晰地将数据点分类的 Hyperplane(超平面)。对于一个线性可分的数据集,SVM 旨在找到最大化 Margin(间隔)的 Hyperplane,即 Hyperplane 与来自任何类别的最近数据点之间的距离。这被称为 Maximum Marginal Hyperplane (MMH),最大间隔超平面。

SVM 中的关键概念:

  • Support Vectors(支持向量):这些是离 Hyperplane 最近的数据点。它们对于定义 Hyperplane 的位置和方向至关重要。如果移动这些点,Hyperplane 也会移动。
  • Hyperplane(超平面):在 2D 空间中,它是一条直线;在 3D 空间中,它是一个平面;在更高维度中,它是一个超平面。它是分隔类别的决策边界。
  • Margin(间隔):Hyperplane 与来自任一类别的最近的支持向量之间的距离。SVMs 试图最大化这个 Margin。

想象一个具有两类点的 2D 散点图。SVM 会试图找到一条直线,用尽可能大的“街道”或“间隙”来分隔这两类点。这条街道的边缘由支持向量定义。

Scikit-Learn 的 SVM 实现支持密集型(NumPy arrays)和稀疏型(SciPy sparse matrices)输入数据。

Scikit-Learn 提供了三个主要的 SVM 分类类:SVC、NuSVC 和 LinearSVC。

SVC (C-Support Vector Classification) 基于 libsvm 库。它可以执行二分类和多分类。对于多分类问题,它默认使用 one-vs-one(一对一)方案(decision_function_shape='ovo'),不过 decision_function_shape='ovr'(one-vs-rest,一对多)也可用,并且通常为了与其他分类器保持一致而更受欢迎。

ParameterDescription
Cfloat, default=1.0。正则化参数。正则化的强度与 C 成反比。必须是严格正值。误差项的惩罚参数 C。
kernel{‘linear’, ‘poly’, ‘rbf’, ‘sigmoid’, ‘precomputed’},default=‘rbf’。指定算法中使用的核函数类型。‘rbf’(径向基函数)是一个常见选择。
degreeint, default=3。多项式核函数(‘poly’)的次数。被所有其他核函数忽略。
gamma{‘scale’, ‘auto’} 或 float, default=‘scale’。‘rbf’、‘poly’ 和 ‘sigmoid’ 核函数的系数。如果 gamma='scale'(默认),则使用 1 / (n_features * X.var())。如果 'auto',则使用 1 / n_features。
coef0float, default=0.0。核函数中的独立项。仅在 ‘poly’ 和 ‘sigmoid’ 中有意义。
shrinkingbool, default=True。是否使用缩减启发式算法,这可以加快训练速度。
probabilitybool, default=False。是否启用概率估计。必须在调用 fit 之前启用,并且会减慢训练速度。
tolfloat, default=1e-3。停止标准的容差。
cache_sizefloat, default=200.0。指定核函数缓存的大小(以 MB 为单位)。
class_weightdict 或 ‘balanced’,default=None。将类别 i 的参数 C 设置为 class_weight[i]*C。如果为 ‘balanced’,权重与类别频率成反比。
verbosebool, default=False。启用详细输出。
max_iterint, default=-1。求解器内部迭代的硬限制,或 -1 表示无限制。
decision_function_shape{‘ovo’, ‘ovr’},default=‘ovr’。是返回形状为 (n_samples, n_classes) 的一对多 (‘ovr’) 决策函数(与所有其他分类器相同),还是返回形状为 (n_samples, n_classes * (n_classes - 1) / 2) 的原始一对一 (‘ovo’) 决策函数。
break_tiesbool, default=False。如果为 True,predict 将根据 decision_function 的置信度值来打破平局;否则返回平局类别中的第一个类别。
random_stateint, RandomState 实例或 None,default=None。当 probability=True 时,控制用于打乱数据以进行概率估计的伪随机性。
AttributeDescription
support_shape 为 (n_SV,) 的 array-like。Support vectors(支持向量)的索引。
support_vectors_shape 为 (n_SV, n_features) 的 array-like。Support vectors(支持向量)。
n_support_shape 为 (n_classes,),dtype 为 int32 的 array-like。每个类别的 Support vectors(支持向量)数量。
dual_coef_array, shape 为 (n_classes-1, n_SV)。决策函数中 Support vectors(支持向量)的系数(用于二分类和多分类 ovo)。
coef_array, shape 为 (n_classes * (n_classes-1) / 2, n_features) 或 (1, n_features) 用于二分类。分配给特征的权重(原始问题中的系数)。仅在 kernel='linear' 时可用。
intercept_array, shape 为 (n_classes * (n_classes-1) / 2,) 或 (1,) 用于二分类。决策函数中的常数项/截距。
classes_shape 为 (n_classes,) 的 array。唯一的类别标签。
import numpy as np
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
# Sample data (linearly separable for simplicity with linear kernel)
X = np.array([[-1, -1], [-2, -1], [1, 1], [2, 1], [-1, 2], [2, -2]])
y = np.array([1, 1, 2, 2, 1, 2]) # Two classes
# It's good practice to scale data for SVMs
svc_pipeline = make_pipeline(StandardScaler(),
SVC(kernel='linear', C=1.0, random_state=42))
svc_pipeline.fit(X, y)
# Making a prediction
new_point = np.array([[-0.8, -0.8]])
prediction = svc_pipeline.predict(new_point)
print(f"Prediction for {new_point}: {prediction}")
# Accessing attributes from the SVC model within the pipeline
svc_model = svc_pipeline.named_steps['svc']
print(f"Support vectors:\n{svc_model.support_vectors_}")
if svc_model.kernel == 'linear':
print(f"Coefficients (w): {svc_model.coef_}")
print(f"Intercept (b): {svc_model.intercept_}")
Prediction for [[-0.8 -0.8]]: [1]
Support vectors:
[[-1. -1.]
[-1. 2.]
[ 1. 1.]]
Coefficients (w): [[ 0.72239822 -0.18059955]]
Intercept (b): [0.03611991]

NuSVC (Nu-Support Vector Classification) 类似于 SVC,但使用参数 nu 来控制支持向量的数量和训练误差。nu 是训练误差比例的上界和支持向量比例的下界。它应该在 (0, 1] 的区间内。

nu: float, default=0.5。替代 SVC 的 C 参数用于控制模型复杂度。

from sklearn.svm import NuSVC
# (Using X, y, StandardScaler, make_pipeline from SVC example)
nusvc_pipeline = make_pipeline(StandardScaler(),
NuSVC(kernel='rbf', nu=0.5, gamma='scale', random_state=42))
nusvc_pipeline.fit(X, y)
prediction_nusvc = nusvc_pipeline.predict(new_point) # Using new_point from SVC example
print(f"NuSVC Prediction for {new_point}: {prediction_nusvc}")
nusvc_model = nusvc_pipeline.named_steps['nusvc']
print(f"Number of support vectors for each class in NuSVC: {nusvc_model.n_support_}")
NuSVC Prediction for [[-0.8 -0.8]]: [1]
Number of support vectors for each class in NuSVC: [3 3]

LinearSVC (Linear Support Vector Classification) 是专门用于线性核函数的 SVM 实现。它基于 liblinear 而不是 libsvm(SVC 和 NuSVC 使用的是 libsvm)。LinearSVC 对于大型数据集通常更快,并且在选择惩罚项(L1/L2)和损失函数(‘hinge’, ‘squared_hinge’)方面提供了更大的灵活性。它没有 kernel 参数(因为它总是线性的)。

与 SVC 的区别:

  • penalty: {‘l1’, ‘l2’},default=‘l2’。指定惩罚中使用的范数。
  • loss: {‘hinge’, ‘squared_hinge’},default=‘squared_hinge’。指定损失函数。‘hinge’ 是标准的 SVM 损失。
  • dual: bool 或 ‘auto’,default=‘auto’。选择算法是求解对偶优化问题还是原始优化问题。当 n_samples > n_features 时优先使用 dual=False。‘auto’ 会自动进行选择。
  • 没有 kernel、gamma、degree、coef0、shrinking、probability 参数。
  • 缺少 support_、support_vectors_、n_support_、dual_coef_ 属性。
from sklearn.svm import LinearSVC
from sklearn.datasets import make_classification
# (StandardScaler, make_pipeline already imported)
# Generate a larger dataset for LinearSVC demonstration
X_lin, y_lin = make_classification(n_samples=200, n_features=4, random_state=0)
linearsvc_pipeline = make_pipeline(StandardScaler(),
LinearSVC(penalty='l2', loss='squared_hinge', dual='auto',
random_state=0, tol=1e-4, C=1.0))
linearsvc_pipeline.fit(X_lin, y_lin)
# Making a prediction on a sample point
sample_lin_point = X_lin[0].reshape(1, -1)
prediction_linearsvc = linearsvc_pipeline.predict(sample_lin_point)
print(f"LinearSVC Prediction for {sample_lin_point[0,:2]}...: {prediction_linearsvc}")
linearsvc_model = linearsvc_pipeline.named_steps['linearsvc']
print(f"LinearSVC Coefficients: {linearsvc_model.coef_}")
print(f"LinearSVC Intercept: {linearsvc_model.intercept_}")
LinearSVC Prediction for [0.16762098 0.20222924]...: [1]
LinearSVC Coefficients: [[ 0.15689569 -0.24839626 0.75454966 0.41789389]]
LinearSVC Intercept: [0.16693748]

SVM 的概念可以扩展到解决回归问题。这被称为 Support Vector Regression (SVR),支持向量回归。在 SVR 中,目标是找到一个函数(Hyperplane),使得尽可能多的训练点与 y 的偏差不超过一个 Margin epsilon,同时该函数尽可能地平坦。在预测值周围 epsilon 范围内的点不会计入损失。

Scikit-Learn 提供了 SVR、NuSVR 和 LinearSVR。

Epsilon-Support Vector Regression,基于 libsvm。主要参数是 C(正则化)和 epsilon(定义不产生误差惩罚的容差范围)。

epsilon: float, default=0.1。指定了在训练损失函数中,预测值与实际值之间距离在 epsilon 范围内不会产生惩罚的 epsilon 范围。

from sklearn.svm import SVR
from sklearn.datasets import make_regression
# (StandardScaler, make_pipeline, np already imported)
X_reg, y_reg = make_regression(n_samples=100, n_features=1, random_state=42, noise=5)
svr_pipeline = make_pipeline(StandardScaler(),
SVR(kernel='rbf', C=100, gamma=0.1, epsilon=0.1))
svr_pipeline.fit(X_reg, y_reg)
sample_reg_point = np.array([[1.5]]) # Example new data point
prediction_svr = svr_pipeline.predict(sample_reg_point)
print(f"SVR Prediction for {sample_reg_point}: {prediction_svr}")
# SVR attributes like support_vectors_ are also available
SVR Prediction for [[1.5]]: [66.9114178]

Nu-Support Vector Regression。类似于 SVR,但使用 nu 来控制支持向量的数量,而不是使用 epsilon 来控制不敏感区域。

from sklearn.svm import NuSVR
# (Using X_reg, y_reg, StandardScaler, make_pipeline, np from SVR example)
nusvr_pipeline = make_pipeline(StandardScaler(),
NuSVR(kernel='rbf', C=100, nu=0.1, gamma='scale'))
nusvr_pipeline.fit(X_reg, y_reg)
prediction_nusvr = nusvr_pipeline.predict(sample_reg_point)
print(f"NuSVR Prediction for {sample_reg_point}: {prediction_nusvr}")
NuSVR Prediction for [[1.5]]: [66.6761768]

Linear Support Vector Regression,基于 liblinear。对于大型数据集上的线性 SVR 更快。主要参数是 epsilon。loss 可以是 'epsilon_insensitive' 或 'squared_epsilon_insensitive'。

from sklearn.svm import LinearSVR
# (Using X_reg, y_reg, StandardScaler, make_pipeline, np from SVR example)
linearsvr_pipeline = make_pipeline(StandardScaler(),
LinearSVR(epsilon=0.0, C=1.0, loss='squared_epsilon_insensitive',
dual='auto', random_state=0, tol=1e-4))
linearsvr_pipeline.fit(X_reg, y_reg)
prediction_linearsvr = linearsvr_pipeline.predict(sample_reg_point)
print(f"LinearSVR Prediction for {sample_reg_point}: {prediction_linearsvr}")
linearsvr_model = linearsvr_pipeline.named_steps['linearsvr']
print(f"LinearSVR Coefficients: {linearsvr_model.coef_}")
print(f"LinearSVR Intercept: {linearsvr_model.intercept_}")
LinearSVR Prediction for [[1.5]]: [67.53298058]
LinearSVR Coefficients: [45.66113378]
LinearSVR Intercept: [-0.95663479]

SVMs 是强大的模型。选择正确的 Kernel(核函数)和参数(如 C、gamma、nu、epsilon)至关重要且通常需要实验(例如,使用 GridSearchCV)。强烈建议进行特征缩放。更多信息,请访问 Scikit-Learn SVM 文档。