Skip to content

使用 Python 实现自编码器

使用 Python 和 Keras 实现一个基本的 Autoencoder

Section titled “使用 Python 和 Keras 实现一个基本的 Autoencoder”

Autoencoder(自编码器)是一种用于无监督学习的人工神经网络 (ANN),用于学习高效的数据编码(表示)。它们通过学习将输入数据压缩到较低维度的“隐空间 (latent space)”,然后从这个压缩表示中重建原始输入来实现这一目标。它们已成为机器学习中用于降维、特征学习和异常检测等任务的重要工具。本章提供了一个分步指南,介绍如何使用 TensorFlow/Keras 在 Python 中实现一个简单的 Autoencoder,并以 MNIST 手写数字数据集作为示例。

我们将涵盖必要的设置、数据预处理、模型构建、训练和结果可视化。本示例重点介绍一个基本的全连接层 (Dense) Autoencoder。

使用 Python 实现 Autoencoder 的分步指南

Section titled “使用 Python 实现 Autoencoder 的分步指南”

让我们探讨使用 Python 和 TensorFlow 中的 Keras API 实现基本 Autoencoder 的步骤。

步骤 1:设置环境 (Setting Up the Environment)

Section titled “步骤 1:设置环境 (Setting Up the Environment)”

开始之前,请确保安装了必要的库。如果没有,您可以使用 pip 进行安装。强烈建议使用虚拟环境 (virtual environment)。

pip install numpy matplotlib tensorflow

步骤 2:导入库 (Importing Libraries)

Section titled “步骤 2:导入库 (Importing Libraries)”

安装完成后,我们需要从这些库中导入必要的模块。

# Import necessary libraries
import numpy as np
import matplotlib.pyplot as plt
from tensorflow.keras.datasets import mnist
from tensorflow.keras.models import Model
from tensorflow.keras.layers import Input, Dense, Flatten, Reshape
from tensorflow.keras.optimizers import Adam
print(f"TensorFlow version: {tf.__version__}") # 可选:检查 TensorFlow 版本

步骤 3:加载和预处理 MNIST 数据集 (Loading and Preprocessing the MNIST Dataset)

Section titled “步骤 3:加载和预处理 MNIST 数据集 (Loading and Preprocessing the MNIST Dataset)”

在这一步中,我们将加载 MNIST 手写数字数据集。像素值将被归一化到 [0, 1] 的范围。由于这是一个简单的全连接层 Autoencoder,我们将把图像展平 (flatten)。

# Load the MNIST dataset (we only need the images, not the labels for unsupervised learning)
# 加载 MNIST 数据集(对于无监督学习,我们只需要图像,不需要标签)
(x_train, _), (x_test, _) = mnist.load_data()
# Normalize the data to the range [0, 1]
# 将数据归一化到 [0, 1] 范围
x_train = x_train.astype('float32') / 255.0
x_test = x_test.astype('float32') / 255.0
# Flatten the 28x28 images into 784-dimensional vectors
# 将 28x28 的图像展平为 784 维向量
x_train_flat = x_train.reshape((len(x_train), np.prod(x_train.shape[1:])))
x_test_flat = x_test.reshape((len(x_test), np.prod(x_test.shape[1:])))
print(f"x_train_flat shape: {x_train_flat.shape}")
print(f"x_test_flat shape: {x_test_flat.shape}")

步骤 4:构建 Autoencoder 模型 (Building the Autoencoder Model)

Section titled “步骤 4:构建 Autoencoder 模型 (Building the Autoencoder Model)”

我们将通过定义编码器 (encoder) 和解码器 (decoder) 部分来构建 Autoencoder 模型。编码器压缩输入,解码器尝试从压缩表示(隐代码,latent code)中重建原始输入。

# Define the dimension of the latent space (compressed representation)
# 定义隐空间(压缩表示)的维度
latent_dim = 64 # Example: compress 784 features to 64
input_dim = x_train_flat.shape[1] # Should be 784 for MNIST
# --- Encoder --- #
input_img = Input(shape=(input_dim,), name='encoder_input')
# Encoded representation
# 编码后的表示
encoded = Dense(latent_dim, activation='relu', name='latent_vector')(input_img)
# --- Decoder --- #
# Decoded representation (reconstruction of the input)
# 解码后的表示(输入的重建)
decoded = Dense(input_dim, activation='sigmoid', name='decoder_output')(encoded) # Sigmoid for pixel values between 0 and 1
# --- Autoencoder Model --- #
# This model maps an input to its reconstruction
# 这个模型将输入映射到其重建
autoencoder = Model(input_img, decoded, name='autoencoder')
# Compile the autoencoder
# 编译 Autoencoder
# Using binary_crossentropy because pixel values are normalized (0-1) and sigmoid activation is used in the output layer.
# This treats each pixel as an independent Bernoulli distribution.
# 使用 binary_crossentropy 是因为像素值已归一化 (0-1),并且输出层使用了 sigmoid 激活函数。
# 这将每个像素视为一个独立的伯努利分布。
autoencoder.compile(optimizer=Adam(learning_rate=0.001), loss='binary_crossentropy')
# Print the summary of the autoencoder model
# 打印 Autoencoder 模型的摘要
autoencoder.summary()

步骤 5:训练 Autoencoder 模型 (Training the Autoencoder Model)

Section titled “步骤 5:训练 Autoencoder 模型 (Training the Autoencoder Model)”

接下来,我们使用训练数据训练 Autoencoder。模型学习从 x_train_flat 本身重建 x_train_flat(即,输入也是目标输出)。

# Train the autoencoder
# 训练 Autoencoder
history = autoencoder.fit(x_train_flat, x_train_flat, # Input and target are the same
epochs=30, # Number of epochs to train
batch_size=256, # Batch size for training
shuffle=True,
validation_data=(x_test_flat, x_test_flat) # 在测试数据上验证
)

步骤 6:可视化原始和重建数据 (Visualizing Original and Reconstructed Data)

Section titled “步骤 6:可视化原始和重建数据 (Visualizing Original and Reconstructed Data)”

在最后一步中,我们将可视化一些原始测试图像以及训练后的 Autoencoder 生成的重建图像,以评估其性能。

# Predict the reconstructed images from the test set
# 从测试集预测重建图像
reconstructed_imgs_flat = autoencoder.predict(x_test_flat)
# Reshape flattened images back to 28x28 to display
# 将展平的图像重新塑形回 28x28 以显示
reconstructed_imgs = reconstructed_imgs_flat.reshape((len(x_test_flat), 28, 28))
original_imgs = x_test # Use the original x_test for comparison, not x_test_flat
# Number of digits to display
# 要显示的数字数量
n = 10
plt.figure(figsize=(20, 4))
for i in range(n):
# Display original image
# 显示原始图像
ax = plt.subplot(2, n, i + 1)
plt.imshow(original_imgs[i], cmap='gray')
plt.title("Original")
plt.axis('off')
# Display reconstructed image
# 显示重建图像
ax = plt.subplot(2, n, i + 1 + n)
plt.imshow(reconstructed_imgs[i], cmap='gray')
plt.title("Reconstructed")
plt.axis('off')
plt.suptitle('Original vs. Reconstructed MNIST Digits', fontsize=16) # 设置图表总标题
plt.tight_layout(rect=[0, 0, 1, 0.96]) # Adjust layout to make space for suptitle
plt.show()

完整的 Python 实现代码 (Complete Python Implementation Code)

Section titled “完整的 Python 实现代码 (Complete Python Implementation Code)”

下面是结合了上述所有步骤的完整 Python 脚本。

import numpy as np
import matplotlib.pyplot as plt
import tensorflow as tf # Import tensorflow
from tensorflow.keras.datasets import mnist
from tensorflow.keras.models import Model
from tensorflow.keras.layers import Input, Dense # Flatten and Reshape not needed for this simple dense AE
from tensorflow.keras.optimizers import Adam
print(f"TensorFlow version: {tf.__version__}")
# Load the MNIST dataset
# 加载 MNIST 数据集
(x_train, _), (x_test, _) = mnist.load_data()
# Normalize the data
# 归一化数据
x_train = x_train.astype('float32') / 255.0
x_test = x_test.astype('float32') / 255.0
# Flatten the 28x28 images into 784-dimensional vectors
# 将 28x28 的图像展平为 784 维向量
x_train_flat = x_train.reshape((len(x_train), np.prod(x_train.shape[1:])))
x_test_flat = x_test.reshape((len(x_test), np.prod(x_test.shape[1:])))
print(f"x_train_flat shape: {x_train_flat.shape}")
print(f"x_test_flat shape: {x_test_flat.shape}")
# Define the dimension of the latent space
# 定义隐空间的维度
latent_dim = 64
input_dim = x_train_flat.shape[1]
# Encoder
# 编码器
input_img = Input(shape=(input_dim,), name='encoder_input')
encoded = Dense(latent_dim, activation='relu', name='latent_vector')(input_img)
# Decoder
# 解码器
decoded = Dense(input_dim, activation='sigmoid', name='decoder_output')(encoded)
# Autoencoder Model
# Autoencoder 模型
autoencoder = Model(input_img, decoded, name='autoencoder')
autoencoder.compile(optimizer=Adam(learning_rate=0.001), loss='binary_crossentropy')
autoencoder.summary()
# Train the autoencoder
# 训练 Autoencoder
print("\nTraining the autoencoder...")
# 训练 Autoencoder...
history = autoencoder.fit(x_train_flat, x_train_flat,
epochs=30, # Reduced epochs for quicker demo, 50 is also fine
# epochs=30, # 减少 epoch 数量以加快演示,50 也可
batch_size=256,
shuffle=True,
validation_data=(x_test_flat, x_test_flat),
verbose=1 # Set to 1 or 2 to see training progress
# verbose=1 # 设置为 1 或 2 查看训练进度
)
print("\nTraining complete. Predicting on test data...")
# 训练完成。正在测试数据上进行预测...
# Predict the reconstructed images from the test set
# 从测试集预测重建图像
reconstructed_imgs_flat = autoencoder.predict(x_test_flat)
# Reshape flattened images back to 28x28 to display
# 将展平的图像重新塑形回 28x28 以显示
reconstructed_imgs = reconstructed_imgs_flat.reshape((len(x_test_flat), 28, 28))
original_imgs = x_test # Use the original x_test for comparison
# Plotting the results
# 绘制结果
print("Plotting original vs. reconstructed images...")
# 正在绘制原始图像与重建图像的对比图...
n = 10
plt.figure(figsize=(20, 4))
for i in range(n):
# Display original image
# 显示原始图像
ax = plt.subplot(2, n, i + 1)
plt.imshow(original_imgs[i], cmap='gray')
plt.title("Original")
plt.axis('off')
# Display reconstructed image
# 显示重建图像
ax = plt.subplot(2, n, i + 1 + n)
plt.imshow(reconstructed_imgs[i], cmap='gray')
plt.title("Reconstructed")
plt.axis('off')
plt.suptitle('Original vs. Reconstructed MNIST Digits by Autoencoder', fontsize=16)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
# Optional: Plot training & validation loss
# 可选:绘制训练损失和验证损失曲线
plt.figure(figsize=(10, 5))
plt.plot(history.history['loss'], label='Training Loss') # 训练损失
plt.plot(history.history['val_loss'], label='Validation Loss') # 验证损失
plt.title('Model Loss During Training') # 训练过程中的模型损失
plt.xlabel('Epoch') # Epoch
plt.ylabel('Loss (Binary Crossentropy)') # 损失 (Binary Crossentropy)
plt.legend()
plt.grid(True)
plt.show()

运行脚本后,它首先会打印 TensorFlow 版本、数据形状和 Autoencoder 模型摘要。然后,它会显示每个 epoch 的训练进度。最后,它会显示两个图表:

  1. 一个图表,顶行显示来自测试集的一些原始 MNIST 数字图像,底行显示训练后的 Autoencoder 重建的对应图像。训练良好的 Autoencoder 将产生与原始图像非常相似的重建图像,这表明模型能够以压缩形式学习和表示数据,尽管由于压缩会丢失一些细节是预期的。

  2. 一个显示训练损失和验证损失随 epoch 变化的曲线图,有助于评估学习过程并检查是否存在过拟合 (overfitting)。

Model: "autoencoder"
_________________________________________________________________
Layer (type) Output Shape Param #
=================================================================
encoder_input (InputLayer) [(None, 784)] 0
latent_vector (Dense) (None, 64) 50240
decoder_output (Dense) (None, 784) 50960
=================================================================
Total params: 101,200
Trainable params: 101,200
Non-trainable params: 0
_________________________________________________________________
Training the autoencoder...
# 训练 Autoencoder...
Epoch 1/30
235/235 [==============================] - 1s 2ms/step - loss: 0.2783 - val_loss: 0.1923
Epoch 2/30
235/235 [==============================] - 0s 2ms/step - loss: 0.1721 - val_loss: 0.1539
...
Epoch 30/30
235/235 [==============================] - 0s 2ms/step - loss: 0.0968 - val_loss: 0.0938
Training complete. Predicting on test data...
# 训练完成。正在测试数据上进行预测...
313/313 [==============================] - 0s 600us/step
Plotting original vs. reconstructed images...
# 正在绘制原始图像与重建图像的对比图...

可视输出随后将显示原始数字与重建数字的比较,然后是损失曲线图。重建的数字应该可以识别,尽管可能比原始图像更模糊一些。

Autoencoder 是强大的无监督学习工具,可应用于各种任务,如降维、特征提取、图像去噪和异常检测。更复杂的架构,如卷积 Autoencoder(用于图像)或变分 Autoencoder(用于生成任务),都建立在这些基本概念之上。

在本章中,我们演示了如何使用 Python 和 Keras 在 MNIST 数据集上实现一个简单的全连接层 Autoencoder。这涉及环境设置、数据预处理、定义模型架构(编码器和解码器)、训练模型以及可视化结果以评估其重建性能。理解这些基本组成部分对于探索更高级的 Autoencoder 变体及其应用至关重要。