Skip to content

SciPy - 聚类

聚类分析(Cluster analysis),或简称聚类(clustering),是将一组对象分组的任务,使得同一组(称为簇,cluster)中的对象彼此之间比与其他组中的对象更相似。K-means 聚类是一种流行的方法,用于将数据集划分为 K 个不同的、不重叠的簇。该算法旨在找到簇中心(centroid),以最小化簇内平方和(inertia)。

K-means 算法迭代执行两个步骤:

  • 分配步骤 (Assignment step): 每个数据点被分配到其质心(centroid)最近的簇(例如,使用欧几里得距离)。
  • 更新步骤 (Update step): 每个簇的质心被重新计算为其分配到的所有数据点的均值。

重复这些步骤,直到质心不再发生显著变化,或达到最大迭代次数。算法收敛后,可以将新的数据点分配到最近的簇质心。SciPy 中的 scipy.cluster.vq 模块(Vector Quantization,向量量化)提供了 K-means 的高效实现。

让我们逐步了解如何使用 scipy.cluster.vq 实现 K-means。

我们将主要使用 kmeans 来查找簇质心,使用 vq 来将观测值分配给簇,以及使用 whiten 进行数据归一化(normalization)。

import numpy as np
from scipy.cluster.vq import kmeans, vq, whiten

为了演示,我们生成一些带有两个不同簇的合成二维(2D)数据。

# Set a seed for reproducibility
np.random.seed(42)
# Generate data for two clusters
# Cluster 1: 50 points around [0, 0]
cluster1_data = np.random.randn(50, 2) + np.array([0, 0])
# Cluster 2: 50 points around [5, 5]
cluster2_data = np.random.randn(50, 2) + np.array([5, 5])
# Combine the data into a single array
# The `features` array will have shape (100, 2)
data_points = np.vstack((cluster1_data, cluster2_data))
print("Shape of the data:", data_points.shape)
print("First 5 data points:\n", data_points[:5])

输出将显示我们数据数组的形状和一部分样本点:

Shape of the data: (100, 2)
First 5 data points:
[[ 0.49671415 -0.1382643 ]
[ 0.64768854 1.52302986]
[-0.23415337 -0.23413696]
[ 1.57921282 0.76743473]
[-0.46947439 0.54256004]]

在应用 K-means 之前,通常最好对数据进行归一化,这在 scipy.cluster.vq 中称为“白化”(whitening)。白化将每个特征(列)重新缩放,使其具有单位方差(unit variance)。这可以确保值较大的特征不会主导距离计算。

# Whiten the data
whitened_data = whiten(data_points)
print("\nFirst 5 whitened data points:\n", whitened_data[:5])
# You can verify that each column now has a standard deviation close to 1
# print("\nStandard deviation of whitened data columns:", np.std(whitened_data, axis=0))

白化会根据每个特征的标准差进行缩放。这是一个常见的预处理步骤。

现在,让我们应用 K-means 算法。我们期望找到两个簇,因此设置 k=2。

# Compute K-Means with k=2 (for two clusters)
# The `kmeans` function returns the centroids and the distortion (average distance)
centroids, distortion = kmeans(whitened_data, 2) # k=2 for two clusters
print("\nCalculated Centroids (for whitened data):\n", centroids)

kmeans 函数在(白化后的)数据上执行 K-means 算法,以找到 k 个簇质心。distortion 值代表观测值与其分配到的质心之间的平均欧几里得距离。

质心的输出将是白化数据空间中的坐标:

Calculated Centroids (for whitened data):
[[-0.02539932 -0.01228082]
[ 1.5908165 1.65607801]]

这些质心对应于算法在归一化特征空间中找到的两个簇的中心。

一旦我们有了质心,就可以使用 vq 函数将每个数据点分配到离它最近的簇。

# Assign each whitened data point to a cluster based on the calculated centroids
# `vq` returns cluster assignments (indices) and distances to centroids
cluster_assignments, distances_to_centroids = vq(whitened_data, centroids)
print("\nCluster assignments for the first 10 data points:", cluster_assignments[:10])
# print("\nDistances for the first 10 data points:", distances_to_centroids[:10])

vq 函数接受数据和质心作为输入,并返回一个数组,其中每个元素是对应数据点所属簇的索引。它还返回每个点到其分配到的质心的距离。

输出将显示前 10 个点的簇索引(0 或 1):

Cluster assignments for the first 10 data points: [0 0 0 0 0 0 0 0 0 0]

在此示例中,由于前 50 个点属于 cluster1_data(中心接近原点),它们大多被分配到一个簇(例如簇 0),而来自 cluster2_data 的后 50 个点(中心接近 [5,5])则被分配到另一个簇(例如簇 1)。确切的分配索引(0 或 1)可能会根据运行情况互换,但分组应该是一致的。您可以检查 cluster_assignments[45:55] 来查看过渡情况。

K-means 聚类广泛应用于各种领域,包括:

  • 客户细分 (Customer Segmentation): 根据购买行为对客户进行分组。
  • 文档聚类 (Document Clustering): 将文本文档按主题组织。
  • 图像分割 (Image Segmentation): 将图像分割成具有相似特征的区域。
  • 异常检测 (Anomaly Detection): 识别不适合任何簇的离群点。

虽然 scipy.cluster.vq 提供了一个基本的 K-means 实现,但像 scikit-learn (sklearn.cluster.KMeans) 这样的库提供了功能更丰富的实现,带有额外的选项和评估指标。