无监督学习:聚类
使用 Python 进行 AI – 无监督学习:聚类
Section titled “使用 Python 进行 AI – 无监督学习:聚类”什么是聚类?
Section titled “什么是聚类?”常用聚类算法
Section titled “常用聚类算法”scikit-learn 库探索几种流行的算法。
K-均值聚类
Section titled “K-均值聚类”指定 K:选择所需的 簇 数量。初始化质心:随机选择 K 个 数据点 作为初始簇质心 (或使用更智能的初始化方法,如 ‘k-means++’)。分配点:将每个 数据点 分配给其质心 最近的簇 (例如,使用欧几里得距离 )。更新质心:将分配给每个 簇 的所有数据点 的均值 重新计算为该簇 的质心 。重复:重复步骤 3 和 4,直到 质心 不再显著移动或达到最大迭代 次数。
pip install scikit-learn matplotlib seaborn numpyimport matplotlib.pyplot as pltimport seaborn as snsimport numpy as npfrom sklearn.cluster import KMeansfrom sklearn.datasets import make_blobs # Updated import
# Generate synthetic 2D data with 4 distinct clustersX, y_true = make_blobs(n_samples=500, centers=4, cluster_std=0.60, random_state=0)
# Visualize the generated dataplt.figure(figsize=(8, 6))plt.scatter(X[:, 0], X[:, 1], s=50, cmap='viridis')plt.title('用于聚类的数据生成')plt.xlabel('特征 1')plt.ylabel('特征 2')plt.show()# Initialize KMeans with K=4 clusters# n_init='auto' is recommended for scikit-learn versions >= 1.4 to use an intelligent default# For older versions, n_init=10 is a common default for 'k-means++'kmeans = KMeans(n_clusters=4, init='k-means++', n_init='auto', random_state=0)
# Fit K-Means to the data and predict cluster labelskmeans.fit(X)y_kmeans = kmeans.predict(X) # Or equivalently: y_kmeans = kmeans.labels_
# Get the cluster centerscenters = kmeans.cluster_centers_
# Visualize the clustering resultsplt.figure(figsize=(10, 7))plt.scatter(X[:, 0], X[:, 1], c=y_kmeans, s=50, cmap='viridis', alpha=0.7, label='数据点')plt.scatter(centers[:, 0], centers[:, 1], c='red', s=200, marker='X', label='质心')plt.title('K-均值聚类结果 (K=4)')plt.xlabel('特征 1')plt.ylabel('特征 2')plt.legend()plt.show()均值漂移聚类
Section titled “均值漂移聚类”初始化窗口:在每个 数据点 周围放置一个窗口 (例如 2D 中的圆形)。计算均值:对于每个 窗口 ,计算其中数据点 的均值 。移动窗口:将 窗口 的中心移动到计算出的均值 位置。重复:重复步骤 2 和 3,直到 窗口 位置收敛 (不再显著移动)。分配聚类: 窗口 收敛到同一位置(模态 )的点被分配到同一簇 。
from sklearn.cluster import MeanShift, estimate_bandwidth
# Generate different sample data for Mean Shiftcenters_ms = [[1, 1], [5, 5], [3, 10]]X_ms, _ = make_blobs(n_samples=300, centers=centers_ms, cluster_std=0.8, random_state=42)
plt.figure(figsize=(8, 6))plt.scatter(X_ms[:, 0], X_ms[:, 1], s=50)plt.title('用于均值漂移的数据生成')plt.show()# Estimate bandwidth (a crucial parameter for Mean Shift)# Bandwidth determines the size of the window.# quantile can be adjusted, 0.2-0.3 often works well.bandwidth = estimate_bandwidth(X_ms, quantile=0.2, n_samples=len(X_ms))
# Initialize MeanShift# If bandwidth is not provided, it's estimated, but providing it can give more control.mean_shift = MeanShift(bandwidth=bandwidth, bin_seeding=True) # bin_seeding can speed up
# Fit Mean Shift to the datamean_shift.fit(X_ms)labels_ms = mean_shift.labels_cluster_centers_ms = mean_shift.cluster_centers_
n_clusters_estimated = len(np.unique(labels_ms))print(f"均值漂移估计的聚类数量:{n_clusters_estimated}")print(f"聚类中心:\n{cluster_centers_ms}")
# Visualize the Mean Shift clustering resultsplt.figure(figsize=(10, 7))colors = ['r', 'g', 'b', 'c', 'm', 'y', 'k']for i in range(len(X_ms)): plt.scatter(X_ms[i, 0], X_ms[i, 1], color=colors[labels_ms[i] % len(colors)], s=50, alpha=0.7)
plt.scatter(cluster_centers_ms[:, 0], cluster_centers_ms[:, 1], marker='X', s=200, linewidths=3, color='black', zorder=10, label='聚类中心')plt.title(f'均值漂移聚类结果(估计聚类数量为 {n_clusters_estimated})')plt.xlabel('特征 1')plt.ylabel('特征 2')plt.legend()plt.show()均值漂移估计的聚类数量: (例如,3)聚类中心:[[x1 y1] [x2 y2] [x3 y3]]衡量聚类性能:轮廓分析
Section titled “衡量聚类性能:轮廓分析”得分接近 +1:样本远离相邻的 簇 (良好的聚类 )。得分接近 0:样本位于或非常接近两个相邻 簇 之间的决策边界 。得分接近 -1:样本可能被分配到错误的 簇 (聚类 效果差)。
s(i) = (b(i) - a(i)) / max(a(i), b(i))
a(i):平均簇内距离 (样本 ‘i’ 到其自身簇 中所有其他点的平均距离)。b(i):平均最近簇距离 (样本 ‘i’ 到次近簇 中所有点的平均距离)。
计算轮廓系数以寻找 K-均值的最优 K 值
Section titled “计算轮廓系数以寻找 K-均值的最优 K 值”from sklearn.metrics import silhouette_score
# Using the data X from the K-Means examplepossible_k_values = range(2, 10) # Test K from 2 to 9silhouette_scores = []
for k_val in possible_k_values: kmeans_eval = KMeans(n_clusters=k_val, init='k-means++', n_init='auto', random_state=0) cluster_labels_eval = kmeans_eval.fit_predict(X)
# Calculate silhouette score score = silhouette_score(X, cluster_labels_eval) silhouette_scores.append(score) print(f"对于 K = {k_val},轮廓系数 = {score:.4f}")
# Plot silhouette scores for different Kplt.figure(figsize=(8, 5))plt.plot(possible_k_values, silhouette_scores, marker='o')plt.title('不同 K 值的轮廓系数 (K-均值)')plt.xlabel('聚类数量 (K)')plt.ylabel('轮廓系数')plt.grid(True)plt.show()
# Find the K with the highest silhouette scoreoptimal_k = possible_k_values[np.argmax(silhouette_scores)]print(f"基于轮廓系数的最佳聚类数量:{optimal_k}")对于 K = 2,轮廓系数 = 0.xxxx对于 K = 3,轮廓系数 = 0.xxxx对于 K = 4,轮廓系数 = 0.xxxx (对于拥有 4 个中心的 make_blobs 示例来说可能最高)...对于 K = 9,轮廓系数 = 0.xxxx
基于轮廓系数的最佳聚类数量:4make_blobs 数据,K=4 应该会产生最高的得分。
sklearn.neighbors.NearestNeighbors:
from sklearn.neighbors import NearestNeighbors
# Sample datasetA = np.array([ [3.1, 2.3], [2.3, 4.2], [3.9, 3.5], [3.7, 6.4], [4.8, 1.9], [8.3, 3.1], [5.2, 7.5], [4.8, 4.7], [3.5, 5.1], [4.4, 2.9]])
# Number of neighbors to findk_neighbors = 3
# Test data point for which we want to find neighborstest_point = np.array([[3.3, 2.9]]) # Must be 2D for kneighbors
# Visualize the input data and the test pointplt.figure(figsize=(8, 6))plt.scatter(A[:, 0], A[:, 1], marker='o', s=100, color='black', label='数据点')plt.scatter(test_point[0, 0], test_point[0, 1], marker='x', s=150, color='red', label='测试点')plt.title('输入数据和测试点')plt.xlabel('特征 1')plt.ylabel('特征 2')plt.legend()plt.grid(True)plt.show()# Initialize and fit the NearestNeighbors model# algorithm='auto' will choose the most appropriate algorithm (e.g., 'ball_tree', 'kd_tree', 'brute')nn_model = NearestNeighbors(n_neighbors=k_neighbors, algorithm='auto')nn_model.fit(A)
# Find the K nearest neighbors for the test pointdistances, indices = nn_model.kneighbors(test_point)
print(f"点 {test_point[0]} 的 {k_neighbors} 个最近邻:")for i in range(k_neighbors): print(f"{i+1}. 点:{A[indices[0][i]]},距离:{distances[0][i]:.4f}")
# Visualize the nearest neighborsplt.figure(figsize=(10, 7))plt.scatter(A[:, 0], A[:, 1], marker='o', s=100, color='black', label='数据点')plt.scatter(test_point[0, 0], test_point[0, 1], marker='x', s=150, color='red', label='测试点')plt.scatter(A[indices[0], 0], A[indices[0], 1], marker='o', s=250, edgecolor='blue', facecolors='none', linewidths=2, label=f'{k_neighbors} 个最近邻')plt.title(f'{k_neighbors}-最近邻')plt.xlabel('特征 1')plt.ylabel('特征 2')plt.legend()plt.grid(True)plt.show()点 [3.3 2.9] 的 3 个最近邻:1. 点:[3.1 2.3],距离:0.63252. 点:[4.4 2.9],距离:1.10003. 点:[3.9 3.5],距离:0.8485(注意:顺序可能会因距离相等的打破规则略有不同,但点应该是正确的。)K-最近邻 (KNN) 分类器
Section titled “K-最近邻 (KNN) 分类器”KNN 分类器概念
Section titled “KNN 分类器概念”选择 K:选择要考虑的 最近邻 的数量(K)。计算距离:对于一个新的、未分类的 数据点 ,计算它与训练数据集 中所有点的距离(例如,欧几里得距离 )。识别邻居:找到与新点最接近的 K 个 训练数据 点。多数投票:将这些 K 个邻居中最频繁出现的类别标签分配给新的 数据点 。(对于回归 ,可以使用邻居值的平均值 )。
示例:用于数字识别的 KNN 分类器
Section titled “示例:用于数字识别的 KNN 分类器”scikit-learn 中的 load_digits 数据集,其中包含手写数字(0-9)的图像。
from sklearn.datasets import load_digitsfrom sklearn.model_selection import train_test_splitfrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.metrics import accuracy_scoreimport matplotlib.pyplot as pltimport numpy as np
# Load the digits datasetdigits = load_digits()X_digits = digits.data # Feature matrix (flattened 8x8 images)y_digits = digits.target # Target labels (0-9)
print(f"数据形状:{X_digits.shape}") # (1797 samples, 64 features)print(f"目标形状:{y_digits.shape}")
# Display a few sample digitsfig, axes = plt.subplots(2, 5, figsize=(10, 5), subplot_kw={'xticks':[], 'yticks':[]})for i, ax in enumerate(axes.flat): ax.imshow(X_digits[i].reshape(8, 8), cmap='binary') ax.set_title(f"标签: {y_digits[i]}")plt.suptitle('数据集中的数字样本')plt.tight_layout(rect=[0, 0, 1, 0.96])plt.show()# Split data into training and testing setsX_train, X_test, y_train, y_test = train_test_split(X_digits, y_digits, test_size=0.3, random_state=42, stratify=y_digits)
print(f"训练样本:{X_train.shape[0]}")print(f"测试样本:{X_test.shape[0]}")
# Initialize KNN classifier (e.g., with K=5 neighbors)knn_classifier = KNeighborsClassifier(n_neighbors=5)
# Train the classifier (for KNN, this mostly involves storing the training data)knn_classifier.fit(X_train, y_train)
# Make predictions on the test sety_pred = knn_classifier.predict(X_test)
# Evaluate the classifieraccuracy = accuracy_score(y_test, y_pred)print(f"KNN 分类器在数字识别上的准确率:{accuracy * 100:.2f}%")
# Let's look at a few test predictionsfig, axes = plt.subplots(3, 4, figsize=(12, 10), subplot_kw={'xticks':[], 'yticks':[]})for i, ax in enumerate(axes.flat): if i < len(X_test): ax.imshow(X_test[i].reshape(8, 8), cmap='binary') true_label = y_test[i] pred_label = y_pred[i] ax.set_title(f"真实值: {true_label}, 预测值: {pred_label}", color='green' if true_label == pred_label else 'red')plt.suptitle('测试预测样本 (KNN 分类器)')plt.tight_layout(rect=[0, 0, 1, 0.96])plt.show()数据形状:(1797, 64)目标形状:(1797,)训练样本:1257测试样本:540
KNN 分类器在数字识别上的准确率:(例如,98.xx % - 会略有不同,但应较高)- Scikit-learn Clustering Documentation: [https://scikit-learn.org/stable/modules/clustering.html]
(Scikit-learn 聚类文档) - Scikit-learn Nearest Neighbors Documentation: [https://scikit-learn.org/stable/modules/neighbors.html]
(Scikit-learn 最近邻文档) - A Visual Guide to K-Means: [https://www.naftaliharris.com/blog/visualizing-k-means-clustering/]
(K-Means 可视化指南)