多层感知器学习
TensorFlow - 使用 Keras 进行多层感知器学习
Section titled “TensorFlow - 使用 Keras 进行多层感知器学习”多层感知器(Multi-Layer Perceptron, MLP)是一种基本类型的人工神经网络。它由多层节点(或称神经元)组成,其中一层中的每个节点都通过一定的权重连接到下一层中的每个节点。通常,一个 MLP 至少包含三层:一个输入层(input layer)、一个或多个隐藏层(hidden layers)和一个输出层(output layer)。
概念上,MLP 的工作方式如下:输入层接收初始数据。每个后续的隐藏层使用激活函数(activation function)转换来自前一层的数据。输出层产生最终结果,例如分类(classification)或回归(regression)值。网络通过调整神经元之间的权重(weight)来学习,通常使用像反向传播(backpropagation)这样的算法。
MLP 通常用于监督学习(supervised learning)任务,例如图像分类(image classification),模型从带标签的数据中学习。反向传播算法是训练 MLP 的基石,它允许网络根据其预测中的误差来调整权重。
让我们使用 MNIST 数据集来实现一个用于图像分类问题的 MLP。我们将使用 TensorFlow 2.x 及其 Keras API,Keras API 提供了一种高级且用户友好的方式来构建和训练神经网络。
# 导入 TensorFlow 和其他必需的库import tensorflow as tffrom tensorflow import kerasfrom tensorflow.keras import layersimport matplotlib.pyplot as pltimport numpy as np
# 参数learning_rate = 0.001training_epochs = 20batch_size = 100display_step = 1
# 网络参数n_hidden_1 = 256 # 第一个隐藏层的特征数量n_hidden_2 = 256 # 第二个隐藏层的特征数量n_input = 784 # MNIST 数据输入(图片形状:28*28 像素)n_classes = 10 # MNIST 总类别数(0-9 数字)
# 使用 Keras API 加载 MNIST 数据(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()
# 预处理数据# 将图片从 28x28 展平为 784x_train = x_train.reshape(-1, n_input).astype('float32') / 255.0x_test = x_test.reshape(-1, n_input).astype('float32') / 255.0
# 将标签转换为独热编码向量y_train = keras.utils.to_categorical(y_train, num_classes=n_classes)y_test = keras.utils.to_categorical(y_test, num_classes=n_classes)
# 使用 Keras Sequential API 构建 MLP 模型model = keras.Sequential([ layers.InputLayer(input_shape=(n_input,)), layers.Dense(n_hidden_1, activation='relu', name='hidden_layer_1'), layers.Dense(n_hidden_2, activation='relu', name='hidden_layer_2'), layers.Dense(n_classes, activation='softmax', name='output_layer')])
# 编译模型# 我们使用 Adam 优化器和 CategoricalCrossentropy 损失函数进行多类别分类。# ReLU 激活函数常用于隐藏层,Softmax 用于分类任务的输出层。optimizer = keras.optimizers.Adam(learning_rate=learning_rate)model.compile(optimizer=optimizer, loss='categorical_crossentropy', metrics=['accuracy'])
# 打印模型概况model.summary()
# 训练模型print("Starting training...")history = model.fit(x_train, y_train, batch_size=batch_size, epochs=training_epochs, validation_data=(x_test, y_test), verbose=1 # 设置为 2 每轮(epoch)一行输出,0 表示静默 )
print("Training phase finished.")
# 绘制训练历史(损失和准确率)epoch_set = range(1, training_epochs + 1)
plt.figure(figsize=(12, 4))
plt.subplot(1, 2, 1)plt.plot(epoch_set, history.history['loss'], 'o-', label='Training Loss')plt.plot(epoch_set, history.history['val_loss'], 'o-', label='Validation Loss')plt.ylabel('Loss')plt.xlabel('Epoch')plt.title('Training and Validation Loss')plt.legend()
plt.subplot(1, 2, 2)plt.plot(epoch_set, history.history['accuracy'], 'o-', label='Training Accuracy')plt.plot(epoch_set, history.history['val_accuracy'], 'o-', label='Validation Accuracy')plt.ylabel('Accuracy')plt.xlabel('Epoch')plt.title('Training and Validation Accuracy')plt.legend()
plt.tight_layout()plt.show()
# 在测试集上评估模型print("\nEvaluating model on test data...")loss, accuracy = model.evaluate(x_test, y_test, verbose=0)print(f"Test Loss: {loss:.4f}")print(f"Test Accuracy: {accuracy:.4f}")
# 有关使用 Keras 进行更详细训练的探索,请参阅:# https://www.tensorflow.org/guide/keras/train_and_evaluate上面的代码定义了一个 MLP,加载并预处理了 MNIST 数据集,然后训练模型。输出将显示模型概况(model summary)、每轮(epoch)的训练进度(损失 loss 和准确率 accuracy),最后是测试准确率(test accuracy)。绘制的图形将可视化显示训练集和验证集在不同轮次(epochs)下的损失和准确率如何变化,这对于诊断过拟合(overfitting)或欠拟合(underfitting)至关重要。
与旧版本 TensorFlow 的主要区别:
- 我们使用
tf.keras.datasets.mnist来加载数据。 - 模型使用
tf.keras.Sequential和tf.keras.layers.Dense构建。 - 激活函数如 ‘relu’(Rectified Linear Unit,修正线性单元)常用于隐藏层,因为它们效率高且能缓解梯度消失问题(vanishing gradient problems)。‘softmax’ 用于多类别分类的输出层。
- 模型使用优化器(optimizer)、损失函数(loss function)和评估指标(metrics)进行编译。
- 训练使用
model.fit()进行,评估使用model.evaluate()。 - TensorFlow 2.x 默认在即时执行模式(eager execution)下运行,使调试更加直观。对于这种编程风格,不再需要 Session(会话)和 placeholder(占位符)。