Skip to content

感知器的隐藏层

TensorFlow - 感知机的隐藏层用于函数逼近

Section titled “TensorFlow - 感知机的隐藏层用于函数逼近”

在本章中,我们将探讨如何使用带有隐藏层的多层感知机(Multi-Layer Perceptron, MLP)来学习逼近一个函数。我们将使用一个带有单个隐藏层的简单神经网络,从一组生成的点 (x, f(x)) 中学习一个余弦函数。这项任务属于回归(regression)的一种形式。

下面的代码演示了如何使用 TensorFlow 2.x 和 Keras API 构建并训练这样一个 MLP。

# 导入必要的模块
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers
import numpy as np
import math
import matplotlib.pyplot as plt
# 设置随机种子以保证结果可复现
np.random.seed(1000)
tf.random.set_seed(1000)
# 定义要学习的函数(带有噪声的余弦函数)
function_to_learn = lambda x: np.cos(x) + 0.1 * np.random.randn(*x.shape)
# 网络和训练参数
num_hidden_neurons = 20 # 增加神经元数量以提高逼近能力
NUM_points = 1000
batch_size = 100
NUM_EPOCHS = 1000 # 减少训练轮数,现代优化器收敛更快
# 生成数据
all_x = np.float32(np.random.uniform(-2 * math.pi, 2 * math.pi, (NUM_points, 1)))
np.random.shuffle(all_x)
# 将数据分割为训练集和验证集
train_size = int(0.9 * NUM_points)
x_training = all_x[:train_size]
y_training = function_to_learn(x_training)
x_validation = all_x[train_size:]
y_validation = function_to_learn(x_validation)
# 绘制生成的数据
plt.figure(1)
plt.scatter(x_training, y_training, c='blue', label='Training Data', s=10)
plt.scatter(x_validation, y_validation, c='pink', label='Validation Data', s=10)
plt.legend()
plt.title('Generated Data for Function Approximation')
plt.xlabel('x')
plt.ylabel('f(x)')
plt.show()
# 使用 Keras Sequential API 构建模型
model = keras.Sequential([
layers.InputLayer(input_shape=(1,)), # 输入层,用于单个特征 x
layers.Dense(num_hidden_neurons, activation='tanh', name='hidden_layer'), # 使用 tanh 激活函数以实现更平滑的逼近
layers.Dense(1, name='output_layer') # 输出层,1 个神经元输出 f(x),默认使用线性激活
])
# 编译模型
# 对于回归任务,均方误差(Mean Squared Error, MSE)是常用的损失函数。
optimizer = keras.optimizers.Adam(learning_rate=0.01) # Adam 优化器
model.compile(optimizer=optimizer, loss='mean_squared_error')
# 打印模型概要
model.summary()
# 训练模型
print("Starting training...")
history = model.fit(x_training, y_training,
batch_size=batch_size,
epochs=NUM_EPOCHS,
validation_data=(x_validation, y_validation),
verbose=0 # 设置为 1 或 2 可查看每轮训练进度
)
print("Training finished.")
# 绘制训练损失和验证损失
plt.figure(2)
plt.plot(history.history['loss'], label='Training Loss')
plt.plot(history.history['val_loss'], label='Validation Loss')
plt.xlabel('Epochs')
plt.ylabel('Cost (MSE)')
plt.title('MLP Function Approximation Training Cost')
plt.legend()
plt.show()
# 在验证集上进行预测,并与实际值进行比较
predictions = model.predict(x_validation)
plt.figure(3)
plt.scatter(x_validation, y_validation, c='pink', label='Actual Validation Data', s=10)
plt.scatter(x_validation, predictions, c='green', label='Model Predictions', s=10)
# 为了更清晰地展示学习到的函数,如果 x_validation 未排序,则对其进行排序
# 并在密集范围的 x 值上对预测结果绘制曲线而非散点。
x_dense = np.linspace(-2 * math.pi, 2 * math.pi, 200).reshape(-1, 1)
y_dense_pred = model.predict(x_dense)
plt.plot(x_dense, y_dense_pred, color='red', label='Learned Function Approximation')
plt.legend()
plt.title('Function Approximation: Actual vs. Predicted')
plt.xlabel('x')
plt.ylabel('f(x)')
plt.show()
# 更多关于使用 Keras 进行回归的信息:
# https://www.tensorflow.org/tutorials/keras/regression

脚本将生成三个图表:

  1. 生成的数据: 此图表展示了带有噪声的余弦函数产生的训练数据点(蓝色)和验证数据点(粉色)的散点图。这可视化了我们的模型将尝试学习的数据。

  2. 训练成本: 此图表显示了训练集和验证集随训练轮数(epochs)变化的均方误差(Mean Squared Error, MSE)。理想情况下,两个损失(loss)都应降低并收敛,这表明模型正在有效学习。训练损失和验证损失之间差距过大可能表明模型过拟合(overfitting)。

  3. 函数逼近: 此图表叠加显示了实际验证数据点(粉色)、模型对这些点的预测结果(绿色),以及一条光滑曲线(红色),该曲线表示模型在密集范围的 x 值上学习到的函数。红色曲线与粉色数据点总体趋势的紧密匹配表明模型成功地进行了逼近。

带有非线性激活函数(本例中使用 tanh,或 ReLU、sigmoid 等)的隐藏层至关重要。它使 MLP 能够学习数据中复杂的非线性关系,这是简单的线性模型无法做到的。