Keras - Convolution Neural Network
Keras - 卷积神经网络 (CNN)
Section titled “Keras - 卷积神经网络 (CNN)”卷积神经网络(CNNs 或 ConvNets)是一类深度神经网络,特别适用于处理网格状数据,最显著的是图像。它们利用共享权重(使用卷积过滤器)和空间层次结构的原理,自动从输入数据中学习特征,因此在图像分类、目标检测和图像分割等任务中非常有效。
让我们将之前使用 MLP 解决的 MNIST 手写数字分类问题,改编为使用 tf.keras 构建的 CNN。CNNs 通常在图像任务上优于 MLPs,因为它们明确地建模了空间关系。
MNIST 的 CNN 概念架构:
输入图像 (28x28x1) -> Conv2D 层(特征提取)-> MaxPooling2D(下采样)-> Conv2D 层 -> MaxPooling2D -> Flatten 层 -> Dense 层(分类头)-> 输出层 (softmax)。
模型规格:
- 数据集: MNIST(手写数字)。
- 预处理: 重塑图像以包含通道维度(例如,28x28x1),归一化像素值。
- 输入层: 由第一个
Conv2D层的input_shape隐式定义。 - 卷积层 1: 带有 32 个过滤器、卷积核大小为 (3,3)、ReLU 激活的
Conv2D。 - 卷积层 2: 带有 64 个过滤器、卷积核大小为 (3,3)、ReLU 激活的
Conv2D。 - 池化层: 在第二个 Conv 层后使用池化窗口大小为 (2,2) 的
MaxPooling2D。 - 正则化: 在池化后使用
Dropout。 - 展平层:
Flatten层将 2D 特征图转换为 1D 向量,供分类器使用。 - Dense 分类头:
Dense层(例如,128 个单元,ReLU 激活),后接Dropout。 - 输出层: 带有 10 个单元(对应每个数字)和
softmax激活的Dense层。 - 编译:
categorical_crossentropy或sparse_categorical_crossentropy损失,Adam优化器(或像原始示例中的 Adadelta),accuracy指标。
步骤 1:导入必要的模块
Section titled “步骤 1:导入必要的模块”import tensorflow as tffrom tensorflow.keras.datasets import mnistfrom tensorflow.keras.models import Sequentialfrom tensorflow.keras.layers import Dense, Dropout, Flattenfrom tensorflow.keras.layers import Conv2D, MaxPooling2Dfrom tensorflow.keras import backend as K # 检查图像数据格式import numpy as np步骤 2:定义参数并加载数据
Section titled “步骤 2:定义参数并加载数据”batch_size = 128num_classes = 10epochs = 12 # 训练轮次
# 输入图像尺寸img_rows, img_cols = 28, 28
# 加载数据,并分割训练集和测试集(x_train, y_train), (x_test, y_test) = mnist.load_data()步骤 3:预处理数据
Section titled “步骤 3:预处理数据”CNNs 需要带有通道维度的输入数据。我们还将像素值归一化到 [0, 1] 范围。
# 根据 Keras 后端数据格式 ('channels_first' 或 'channels_last') 确定输入形状if K.image_data_format() == 'channels_first': x_train = x_train.reshape(x_train.shape[0], 1, img_rows, img_cols) x_test = x_test.reshape(x_test.shape[0], 1, img_rows, img_cols) input_shape = (1, img_rows, img_cols)else: # 'channels_last' x_train = x_train.reshape(x_train.shape[0], img_rows, img_cols, 1) x_test = x_test.reshape(x_test.shape[0], img_rows, img_cols, 1) input_shape = (img_rows, img_cols, 1)
# 将数据类型转换为 float32 并归一化x_train = x_train.astype('float32') / 255.0x_test = x_test.astype('float32') / 255.0
print('x_train shape:', x_train.shape)print(x_train.shape[0], 'train samples')print(x_test.shape[0], 'test samples')
# 将类别向量转换为二进制类别矩阵 (独热编码)# 对于 'categorical_crossentropy' 损失是必要的y_train_cat = tf.keras.utils.to_categorical(y_train, num_classes)y_test_cat = tf.keras.utils.to_categorical(y_test, num_classes)
# 如果使用 'sparse_categorical_crossentropy',保持 y_train/y_test 为整数。步骤 4:构建 CNN 模型
Section titled “步骤 4:构建 CNN 模型”model = Sequential([ tf.keras.Input(shape=input_shape), # 定义输入形状 Conv2D(32, kernel_size=(3, 3), activation='relu'), Conv2D(64, (3, 3), activation='relu'), MaxPooling2D(pool_size=(2, 2)), Dropout(0.25), Flatten(), # 将 2D 特征图展平为 1D Dense(128, activation='relu'), Dropout(0.5), Dense(num_classes, activation='softmax')], name="mnist_cnn")
model.summary()步骤 5:编译模型
Section titled “步骤 5:编译模型”配置学习过程。因为标签是独热编码的,我们将使用 categorical cross-entropy。
# 优化器:原始示例使用了 Adadelta(),Adam 也常用。optimizer = tf.keras.optimizers.Adam() # Or tf.keras.optimizers.Adadelta()
model.compile(loss=tf.keras.losses.categorical_crossentropy, optimizer=optimizer, metrics=['accuracy'])步骤 6:训练模型
Section titled “步骤 6:训练模型”将模型拟合到训练数据。
print("\nTraining CNN Model...") # 训练 CNN 模型...history = model.fit(x_train, y_train_cat, # 使用独热编码的标签 batch_size=batch_size, epochs=epochs, verbose=1, validation_data=(x_test, y_test_cat))
print("\nTraining Complete.") # 训练完成。监控训练期间打印的训练和验证准确率/损失。CNNs 通常在 MNIST 上达到比简单 MLPs 更高的准确率。
# 示例输出片段(值取决于优化器、轮次等):# Epoch 1/12# Epoch 2/12# 469/469 [==============================] - 8s 18ms/step - loss: 0.0876 - accuracy: 0.9741 - val_loss: 0.0396 - val_accuracy: 0.9869# ...# Epoch 12/12# 469/469 [==============================] - 9s 18ms/step - loss: 0.0186 - accuracy: 0.9941 - val_loss: 0.0272 - val_accuracy: 0.9917步骤 7:评估模型
Section titled “步骤 7:评估模型”评估在测试集上的最终性能。
print("\nEvaluating Model...") # 评估模型...score = model.evaluate(x_test, y_test_cat, verbose=0)print(f'Test loss: {score[0]:.4f}') # 测试损失:print(f'Test accuracy: {score[1]:.4f}') # 测试准确率:CNNs 通常可以在 MNIST 上相对轻松地达到 >99% 的准确率。
步骤 8:进行预测 (可选)
Section titled “步骤 8:进行预测 (可选)”使用训练好的模型预测测试图像的标签。
# 获取预测结果 (概率)predictions_probabilities = model.predict(x_test)
# 获取预测标签 (最大概率的索引)predicted_labels = np.argmax(predictions_probabilities, axis=1)
# 比较前几个预测结果与实际标签print("\nFirst 5 Predicted labels:", predicted_labels[:5]) # 前 5 个预测标签:print("First 5 Actual labels: ", y_test[:5]) # 前 5 个实际标签: (原始整数标签)这个示例展示了使用 tf.keras 构建、训练和评估用于图像分类的 CNN 的典型工作流程,重点介绍了 Conv2D 和 MaxPooling2D 层的使用。