SciPy - 聚类
SciPy - 聚类分析
Section titled “SciPy - 聚类分析”聚类分析(Cluster analysis),或简称聚类(clustering),是将一组对象分组的任务,使得同一组(称为簇,cluster)中的对象彼此之间比与其他组中的对象更相似。K-means 聚类是一种流行的方法,用于将数据集划分为 K 个不同的、不重叠的簇。该算法旨在找到簇中心(centroid),以最小化簇内平方和(inertia)。
K-means 算法迭代执行两个步骤:
- 分配步骤 (Assignment step): 每个数据点被分配到其质心(centroid)最近的簇(例如,使用欧几里得距离)。
- 更新步骤 (Update step): 每个簇的质心被重新计算为其分配到的所有数据点的均值。
重复这些步骤,直到质心不再发生显著变化,或达到最大迭代次数。算法收敛后,可以将新的数据点分配到最近的簇质心。SciPy 中的 scipy.cluster.vq 模块(Vector Quantization,向量量化)提供了 K-means 的高效实现。
SciPy 中的 K-Means 实现
Section titled “SciPy 中的 K-Means 实现”让我们逐步了解如何使用 scipy.cluster.vq 实现 K-means。
导入所需函数
Section titled “导入所需函数”我们将主要使用 kmeans 来查找簇质心,使用 vq 来将观测值分配给簇,以及使用 whiten 进行数据归一化(normalization)。
import numpy as npfrom scipy.cluster.vq import kmeans, vq, whiten为了演示,我们生成一些带有两个不同簇的合成二维(2D)数据。
# Set a seed for reproducibilitynp.random.seed(42)
# Generate data for two clusters# Cluster 1: 50 points around [0, 0]cluster1_data = np.random.randn(50, 2) + np.array([0, 0])
# Cluster 2: 50 points around [5, 5]cluster2_data = np.random.randn(50, 2) + np.array([5, 5])
# Combine the data into a single array# The `features` array will have shape (100, 2)data_points = np.vstack((cluster1_data, cluster2_data))
print("Shape of the data:", data_points.shape)print("First 5 data points:\n", data_points[:5])输出将显示我们数据数组的形状和一部分样本点:
Shape of the data: (100, 2)First 5 data points: [[ 0.49671415 -0.1382643 ] [ 0.64768854 1.52302986] [-0.23415337 -0.23413696] [ 1.57921282 0.76743473] [-0.46947439 0.54256004]]数据白化(归一化)
Section titled “数据白化(归一化)”在应用 K-means 之前,通常最好对数据进行归一化,这在 scipy.cluster.vq 中称为“白化”(whitening)。白化将每个特征(列)重新缩放,使其具有单位方差(unit variance)。这可以确保值较大的特征不会主导距离计算。
# Whiten the datawhitened_data = whiten(data_points)
print("\nFirst 5 whitened data points:\n", whitened_data[:5])# You can verify that each column now has a standard deviation close to 1# print("\nStandard deviation of whitened data columns:", np.std(whitened_data, axis=0))白化会根据每个特征的标准差进行缩放。这是一个常见的预处理步骤。
使用两个簇计算 K-Means
Section titled “使用两个簇计算 K-Means”现在,让我们应用 K-means 算法。我们期望找到两个簇,因此设置 k=2。
# Compute K-Means with k=2 (for two clusters)# The `kmeans` function returns the centroids and the distortion (average distance)centroids, distortion = kmeans(whitened_data, 2) # k=2 for two clusters
print("\nCalculated Centroids (for whitened data):\n", centroids)kmeans 函数在(白化后的)数据上执行 K-means 算法,以找到 k 个簇质心。distortion 值代表观测值与其分配到的质心之间的平均欧几里得距离。
质心的输出将是白化数据空间中的坐标:
Calculated Centroids (for whitened data): [[-0.02539932 -0.01228082] [ 1.5908165 1.65607801]]这些质心对应于算法在归一化特征空间中找到的两个簇的中心。
将每个观测值分配到簇
Section titled “将每个观测值分配到簇”一旦我们有了质心,就可以使用 vq 函数将每个数据点分配到离它最近的簇。
# Assign each whitened data point to a cluster based on the calculated centroids# `vq` returns cluster assignments (indices) and distances to centroidscluster_assignments, distances_to_centroids = vq(whitened_data, centroids)
print("\nCluster assignments for the first 10 data points:", cluster_assignments[:10])# print("\nDistances for the first 10 data points:", distances_to_centroids[:10])vq 函数接受数据和质心作为输入,并返回一个数组,其中每个元素是对应数据点所属簇的索引。它还返回每个点到其分配到的质心的距离。
输出将显示前 10 个点的簇索引(0 或 1):
Cluster assignments for the first 10 data points: [0 0 0 0 0 0 0 0 0 0]在此示例中,由于前 50 个点属于 cluster1_data(中心接近原点),它们大多被分配到一个簇(例如簇 0),而来自 cluster2_data 的后 50 个点(中心接近 [5,5])则被分配到另一个簇(例如簇 1)。确切的分配索引(0 或 1)可能会根据运行情况互换,但分组应该是一致的。您可以检查 cluster_assignments[45:55] 来查看过渡情况。
K-means 聚类广泛应用于各种领域,包括:
- 客户细分 (Customer Segmentation): 根据购买行为对客户进行分组。
- 文档聚类 (Document Clustering): 将文本文档按主题组织。
- 图像分割 (Image Segmentation): 将图像分割成具有相似特征的区域。
- 异常检测 (Anomaly Detection): 识别不适合任何簇的离群点。
虽然 scipy.cluster.vq 提供了一个基本的 K-means 实现,但像 scikit-learn (sklearn.cluster.KMeans) 这样的库提供了功能更丰富的实现,带有额外的选项和评估指标。