Skip to content

聚类性能评估

评估聚类算法的性能不像评估监督学习模型(如分类器或回归器)那样直接,因为无监督任务通常没有真实标签 (ground truth labels)。然而,Scikit-learn 在 sklearn.metrics 模块中提供了几种指标来评估聚类结果的质量。这些指标大致可分为两类:

  • 外部评估指标 (Extrinsic Measures):需要真实标签。当已知真实聚类分配时使用(例如,用于基准测试或在标签存在但未在聚类期间使用的数据集上)。示例:调整兰德指数 (Adjusted Rand Index)、互信息分数 (Mutual Information scores)。
  • 内部评估指标 (Intrinsic Measures):不需要真实标签。根据聚类的固有属性进行评估,例如紧密度 (cohesion)(簇内点之间的距离)和分离度 (separation)(簇之间的区别)。示例:轮廓系数 (Silhouette Coefficient)。

外部评估指标(需要真实标签)

Section titled “外部评估指标(需要真实标签)”

这些指标将算法的预测聚类分配 (labels_pred) 与真实的类别标签 (labels_true) 进行比较。

调整兰德指数 (Adjusted Rand Index, ARI)

Section titled “调整兰德指数 (Adjusted Rand Index, ARI)”

兰德指数 (Rand Index, RI) 通过考虑所有样本对并计算在预测和真实聚类中都被分配到相同或不同簇中的对数,来衡量两种数据聚类之间的相似性。调整兰德指数 (Adjusted Rand Index, ARI) 是 RI 的一个校正了偶然性的版本。

公式(概念性):ARI = (RI - Expected_RI) / (max(RI) - Expected_RI)。

值范围从 -1 到 1。得分为 1 表示完美一致,0 表示随机标记(与真实标签无关),负值表示比随机更差。由 sklearn.metrics.adjusted_rand_score 提供。

from sklearn.metrics import adjusted_rand_score
labels_true = [0, 0, 0, 1, 1, 1]
labels_pred_good = [0, 0, 0, 1, 1, 1] # 完美匹配
labels_pred_bad = [0, 1, 0, 1, 0, 1] # 匹配不佳
labels_pred_permuted = [1, 1, 1, 0, 0, 0] # 置换后的完美匹配
ari_good = adjusted_rand_score(labels_true, labels_pred_good)
ari_bad = adjusted_rand_score(labels_true, labels_pred_bad)
ari_permuted = adjusted_rand_score(labels_true, labels_pred_permuted)
print(f"ARI (良好): {ari_good:.4f}")
print(f"ARI (不佳): {ari_bad:.4f}")
print(f"ARI (置换良好): {ari_permuted:.4f}")
ARI (良好): 1.0000
ARI (不佳): -0.0714
ARI (置换良好): 1.0000

互信息 (Mutual Information, MI) 衡量两种分配之间的一致性,忽略置换。Scikit-learn 提供几种基于 MI 的分数:

  • 归一化互信息 (Normalized Mutual Information, NMI):sklearn.metrics.normalized_mutual_info_score。MI 由熵归一化 (normalized by entropy)。范围从 0(无互信息)到 1(完美相关)。
  • 调整互信息 (Adjusted Mutual Information, AMI):sklearn.metrics.adjusted_mutual_info_score。校正了偶然性 (adjusted for chance) 的 MI。与 ARI 类似,值越高越好,1 表示完美一致。
from sklearn.metrics import normalized_mutual_info_score, adjusted_mutual_info_score
labels_true = [0, 0, 1, 1, 1, 1]
labels_pred = [0, 0, 2, 2, 3, 3] # 簇标签不同,但结构某种程度上得以保留
nmi_score = normalized_mutual_info_score(labels_true, labels_pred)
ami_score = adjusted_mutual_info_score(labels_true, labels_pred)
print(f"归一化互信息 (NMI): {nmi_score:.4f}")
print(f"调整互信息 (AMI): {ami_score:.4f}")
归一化互信息 (NMI): 0.7612
调整互信息 (AMI): 0.4444

同质性 (Homogeneity)、完整性 (Completeness) 和 V-measure

Section titled “同质性 (Homogeneity)、完整性 (Completeness) 和 V-measure”

这三个指标提供了对聚类质量不同方面的洞察:

  • 同质性 (Homogeneity) (sklearn.metrics.homogeneity_score):每个簇只包含单个类别的成员。得分为 1 表示所有簇都完全同质。
  • 完整性 (Completeness) (sklearn.metrics.completeness_score):给定类别的所有成员都被分配到同一个簇。得分为 1 表示给定类别的所有数据点都落入一个簇中。
  • V-measure (sklearn.metrics.v_measure_score):同质性和完整性的调和平均数 (harmonic mean)。得分为 1 表示完美同质性和完整性。

这些分数的范围是 0 到 1,1 是最优值。

from sklearn.metrics import homogeneity_score, completeness_score, v_measure_score
labels_true = [0, 0, 0, 1, 1, 1]
labels_pred = [0, 0, 0, 1, 2, 2] # 类别 1 被分割到两个簇中
hom_score = homogeneity_score(labels_true, labels_pred)
com_score = completeness_score(labels_true, labels_pred)
vm_score = v_measure_score(labels_true, labels_pred)
print(f"同质性 (Homogeneity): {hom_score:.4f}")
print(f"完整性 (Completeness): {com_score:.4f}")
print(f"V-measure: {vm_score:.4f}")
同质性 (Homogeneity): 0.6667
完整性 (Completeness): 1.0000
V-measure: 0.8000

这里,完整性为 1,因为所有真实的 ‘0’ 都在簇 ‘0’ 中,并且所有真实的 ‘1’ 都被统计在内(尽管被分割)。同质性较低,因为如果我们考虑原始类别,预测中的簇 ‘1’ 和 ‘2’ 并非纯粹来自单一的真实类别。

Fowlkes-Mallows 得分 (Fowlkes-Mallows Score, FMI)

Section titled “Fowlkes-Mallows 得分 (Fowlkes-Mallows Score, FMI)”

Fowlkes-Mallows 得分 (sklearn.metrics.fowlkes_mallows_score) 定义为成对精确率 (pairwise precision) 和召回率 (recall) 的几何平均数 (geometric mean)。它衡量从真实标签和预测标签获得的簇之间的相似性。

FMI = TP / sqrt((TP + FP) * (TP + FN))

其中 TP 是真阳性 (true positives) 的数量(在真实和预测聚类中都在同一簇中的点对)。FP 是假阳性 (false positives),FN 是假阴性 (false negatives)。得分范围从 0 到 1,值越高越好。

from sklearn.metrics import fowlkes_mallows_score
labels_true = [0, 0, 1, 1, 1, 1]
labels_pred = [0, 0, 2, 2, 3, 3]
fmi = fowlkes_mallows_score(labels_true, labels_pred)
print(f"Fowlkes-Mallows 得分: {fmi:.4f}")
Fowlkes-Mallows 得分: 0.6547

内部评估指标(不需要真实标签)

Section titled “内部评估指标(不需要真实标签)”

轮廓系数 (sklearn.metrics.silhouette_score) 衡量一个对象与它自己的簇有多相似(内聚性 cohesion),相比于与其他簇的相似度(分离度 separation)。它是为每个样本计算的。

对于样本 i,令 a(i) 为平均簇内距离 (mean intra-cluster distance)(样本 i 到其自身簇中所有其他点的平均距离),b(i) 为平均最近簇距离 (mean nearest-cluster distance)(样本 i 到次近簇中所有点的平均距离)。

样本 i 的轮廓系数:s(i) = (b(i) - a(i)) / max(a(i), b(i))

一组样本的总体轮廓得分 (Overall Silhouette Score) 是所有样本的 s(i) 的平均值。值范围从 -1 到 1:

  • 接近 +1:样本距离相邻簇很远(聚类良好)。
  • 0:样本位于或非常接近两个相邻簇之间的决策边界。
  • 接近 -1:样本可能被分配到了错误的簇。
from sklearn.metrics import silhouette_score
from sklearn.datasets import load_iris
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
import numpy as np
# 加载 iris 数据集
X, y = load_iris(return_X_y=True)
# 为 K-Means 缩放数据
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
# 应用 K-Means(以 Iris 数据集为例,使用 3 个簇)
kmeans_model = KMeans(n_clusters=3, random_state=42, n_init='auto').fit(X_scaled)
cluster_labels = kmeans_model.labels_
# 计算轮廓得分
# 需要数据 X 和预测标签
sil_score = silhouette_score(X_scaled, cluster_labels, metric='euclidean')
print(f"K-Means (k=3) 在 Iris 数据集上的轮廓得分: {sil_score:.4f}")
K-Means (k=3) 在 Iris 数据集上的轮廓得分: 0.4599

轮廓得分常用于帮助选择最优的簇数量 (k),方法是尝试不同的 k 值并选择使得分最大化的那个值。

Scikit-learn 还提供其他内部评估指标,例如:

  • Calinski-Harabasz 指数 (Calinski-Harabasz Index) (sklearn.metrics.calinski_harabasz_score):也称为方差比准则 (Variance Ratio Criterion)。值越高表示簇的定义越好。
  • Davies-Bouldin 指数 (Davies-Bouldin Index) (sklearn.metrics.davies_bouldin_score):衡量每个簇与其最相似簇的平均相似度比率。值越低表示聚类效果越好(越接近 0)。

如果已知真实类别标签,混淆矩阵 (sklearn.metrics.cluster.contingency_matrix) 可以提供洞见。它显示了每个真实类别 / 预测簇对的交集基数 (intersection cardinality)。这类似于混淆矩阵 (confusion matrix),但用于聚类上下文。

from sklearn.metrics.cluster import contingency_matrix
import numpy as np
labels_true = np.array(["cat", "cat", "cat", "dog", "dog", "bird"])
labels_pred = np.array([0, 0, 1, 1, 2, 2]) # 0: 簇 A, 1: 簇 B, 2: 簇 C
# 假设唯一的真实标签是 ['bird', 'cat', 'dog'],预测标签是 [0, 1, 2]
# (默认按字母/数字排序)
cont_matrix = contingency_matrix(labels_true, labels_pred)
print("混淆矩阵(行=真实,列=预测):")
print(cont_matrix)
# 为了更好地理解行/列:
# unique_true_labels = np.unique(labels_true)
# unique_pred_labels = np.unique(labels_pred)
# print(f"真实标签(行): {unique_true_labels}")
# print(f"预测标签(列): {unique_pred_labels}")
混淆矩阵(行=真实,列=预测):
[[0 0 1] # <-- 真实 'bird'
[2 1 0] # <-- 真实 'cat'
[0 1 1]] # <-- 真实 'dog'

例如,第一行 [0 0 1] 表示:对于真实标签 ‘bird’(第一个唯一的真实标签),有 0 个样本被放入预测簇 0,0 个放入簇 1,1 个放入簇 2。第二行 [2 1 0] 表示:对于真实标签 ‘cat’,有 2 个样本被放入簇 0,1 个放入簇 1,0 个放入簇 2。

有关这些指标及其用法的更多详细信息,请参阅 Scikit-learn 文档:https://scikit-learn.org/stable/modules/clustering.html#clustering-performance-evaluation