Skip to content

从头训练一个卷积网络

在本章中,我们将重点介绍如何使用 PyTorch 的 nn.Module 从零开始创建和训练一个简单的前馈神经网络(feed-forward neural network)。这包括定义网络架构、前向传播,然后实现一个训练循环。

我们将创建一个继承自 torch.nn.Module 的类。在类内部,我们定义网络的各个层(layers)。对于一个简单的网络,这可能包括线性层(linear layers)和激活函数(activation functions)。

import torch
import torch.nn as nn
import torch.optim as optim
class SimpleNeuralNetwork(nn.Module):
def __init__(self, input_size, hidden_size, output_size):
super(SimpleNeuralNetwork, self).__init__()
self.input_size = input_size
self.hidden_size = hidden_size
self.output_size = output_size
# Define layers
# 定义层
self.fc1 = nn.Linear(self.input_size, self.hidden_size) # Input layer to hidden layer
# 输入层到隐藏层
self.relu = nn.ReLU() # Activation function for hidden layer
# 隐藏层的激活函数
self.fc2 = nn.Linear(self.hidden_size, self.output_size) # Hidden layer to output layer
# 隐藏层到输出层
self.sigmoid = nn.Sigmoid() # Activation function for output layer (e.g., for binary classification)
# 输出层的激活函数(例如,用于二分类)
def forward(self, x):
# Define the forward pass
# 定义前向传播
out = self.fc1(x)
out = self.relu(out)
out = self.fc2(out)
out = self.sigmoid(out) # Use sigmoid if output is a probability (0-1)
# 如果输出是概率 (0-1),则使用 sigmoid
return out
# Example instantiation:
# 示例实例化:
# input_features = 10
# 输入特征数 = 10
# hidden_units = 20
# 隐藏单元数 = 20
# output_classes = 1 (for binary classification)
# 输出类别数 = 1 (用于二分类)
# model = SimpleNeuralNetwork(input_features, hidden_units, output_classes)

与旧方法的关键区别:

  • 我们使用 nn.Linear 表示全连接层(fully connected layers),它在内部处理权重(weight)和偏置(bias)的初始化。
  • 我们使用 torch.nn 中的标准激活函数,如 nn.ReLU 和 nn.Sigmoid。
  • forward 方法定义了数据如何流经网络(即前向传播)。

步骤 2:准备数据、损失函数和优化器

Section titled “步骤 2:准备数据、损失函数和优化器”

在训练之前,你需要数据、用于衡量误差的损失函数(loss function),以及用于更新模型权重的优化器(optimizer)。

# Example: Dummy data for illustration
# 示例:用于说明的模拟数据
# Assume X_train is your input data and y_train are your labels
# 假设 X_train 是输入数据,y_train 是标签
# For a network with input_size=2, output_size=1:
# 对于 input_size=2, output_size=1 的网络:
X_train = torch.randn(100, 2) # 100 samples, 2 features each
# 100 个样本,每个样本有 2 个特征
y_train = torch.randint(0, 2, (100, 1)).float() # 100 binary labels
# 100 个二分类标签
# Instantiate the model
# 实例化模型
input_dim = 2
hidden_dim = 3
output_dim = 1
model = SimpleNeuralNetwork(input_dim, hidden_dim, output_dim)
# Loss Function
# 损失函数
# For binary classification with sigmoid output, BCELoss is common.
# 对于使用 sigmoid 输出的二分类,BCELoss 是常见的选择。
# If it were a regression task, MSELoss might be used.
# 如果是回归任务,可能会使用 MSELoss。
criterion = nn.BCELoss()
# Optimizer
# 优化器
# Adam and SGD are popular choices.
# Adam 和 SGD 是流行的选择。
learning_rate = 0.01
optimizer = optim.Adam(model.parameters(), lr=learning_rate)
print("Model Architecture:")
print(model)

PyTorch 的 autograd 系统会自动计算梯度(gradients)。我们不再需要手动实现反向传播逻辑或导数函数,如 sigmoidPrime。

训练循环(training loop)会多次迭代遍历数据(每个迭代称为一个 epoch),执行前向传播(forward pass)和反向传播(backward pass)。

num_epochs = 100
for epoch in range(num_epochs):
# Forward pass: Compute predicted y by passing X to the model
# 前向传播:将 X 输入模型计算预测的 y 值
outputs = model(X_train)
loss = criterion(outputs, y_train)
# Backward pass and optimize
# 反向传播并优化
optimizer.zero_grad() # Zero the gradients before running the backward pass.
# 在运行反向传播之前将梯度清零。
loss.backward() # Perform backpropagation: compute gradients of the loss w.r.t. model parameters.
# 执行反向传播:计算损失相对于模型参数的梯度。
optimizer.step() # Calling step() causes the optimizer to update the parameters.
# 调用 step() 会使优化器更新参数。
if (epoch+1) % 10 == 0:
print(f'Epoch [{epoch+1}/{num_epochs}], Loss: {loss.item():.4f}')
print("\nTraining finished.")

训练完成后,你可以使用模型进行预测并保存其状态。

# Prediction
# 预测
def predict(model, X_new):
model.eval() # Set the model to evaluation mode (important for layers like Dropout, BatchNorm)
# 将模型设置为评估模式(对于 Dropout、BatchNorm 等层很重要)
with torch.no_grad(): # Disable gradient calculations for inference
# 推断时禁用梯度计算
predictions = model(X_new)
predicted_classes = (predictions > 0.5).float() # Assuming binary classification with 0.5 threshold
# 假设二分类,阈值为 0.5
return predicted_classes
# Example prediction with new data
# 使用新数据进行预测示例
X_test = torch.randn(5, 2) # 5 new samples
# 5 个新样本
predictions = predict(model, X_test)
print("\nPredicted data based on trained weights: ")
# 基于训练权重预测的数据:
print("Input (scaled): \n" + str(X_test))
# 输入 (缩放后):
print("Output (predicted classes): \n" + str(predictions))
# 输出 (预测类别):
# Saving and Loading Model Weights
# 保存和加载模型权重
MODEL_PATH = "simple_nn.pth"
# Save model state_dict (recommended)
# 保存模型 state_dict (推荐)
torch.save(model.state_dict(), MODEL_PATH)
print(f'\nModel state_dict saved to {MODEL_PATH}')
# To load the model weights:
# 加载模型权重:
# First, instantiate the model again with the same architecture
# 首先,用相同的架构再次实例化模型
loaded_model = SimpleNeuralNetwork(input_dim, hidden_dim, output_dim)
loaded_model.load_state_dict(torch.load(MODEL_PATH))
loaded_model.eval() # Set to evaluation mode after loading
# 加载后设置为评估模式
print('Model loaded successfully.')
# You can also save the entire model (less flexible):
# 也可以保存整个模型(灵活性较低):
# torch.save(model, "entire_model.pth")
# loaded_model = torch.load("entire_model.pth")

在推断(inference)或验证(validation)期间使用 model.eval() 非常重要,以确保 Dropout 或 Batch Normalization 等层的行为正确。torch.no_grad() 会禁用梯度计算,从而节省内存并加快推断速度。