Apache MXNet - Python API gluon
Apache MXNet - Gluon API 深度探索
Section titled “Apache MXNet - Gluon API 深度探索”MXNet Gluon API 提供了一个用户友好的、命令式(imperative)接口,用于设计、训练和部署深度学习模型。它通过其混合(hybrid)功能,将动态框架的易用性与符号式执行(symbolic execution)的性能结合起来。
核心 Gluon 模块
Section titled “核心 Gluon 模块”本章将探索 mxnet.gluon 包中的关键模块。
gluon.nn:神经网络层
Section titled “gluon.nn:神经网络层”mxnet.gluon.nn 模块是构建模型的核心,提供了丰富的预定义神经网络层(Layer)集合。
关键层(Layer)和块(Block)
Section titled “关键层(Layer)和块(Block)”gluon.nn 中的一些重要类(class):
| 类名 (Class Name) | 描述 (Description) |
|---|---|
Block([prefix, params]) | Gluon 中所有神经网络层和模型的基类。通过继承 Block 可以构建自定义模型。 |
HybridBlock([prefix, params]) | Block 的子类,支持命令式执行和符号图编译(混合化,hybridization),用于性能优化。 |
Sequential([prefix, params]) | 用于按顺序堆叠层(Layer)的容器。HybridSequential 是其对应的混合版本。 |
Dense(units, activation=None, use_bias=True, ...) | 标准的ᵉ全连接层(fully connected layer,或称密集层,dense layer)。 |
Conv1D, Conv2D, Conv3D | 分别对应 1D、2D 和 3D 卷积层(convolution layer)。参数包括 channels(通道数)、kernel_size(卷积核大小)、strides(步长)、padding(填充)、activation(激活函数)。 |
Conv*DTranspose | 转置卷积层(Transposed convolution layer,常被称为反卷积层,deconvolution layer),用于上采样(upsampling)。 |
Activation(activation_type) | 应用指定的激活函数(activation function),例如 ‘relu’、‘sigmoid’、‘tanh’。 |
BatchNorm([axis, momentum, epsilon, ...]) | 批量归一化层(Batch normalization layer),用于稳定训练并提高泛化能力。 |
Dropout(rate, axes=()) | 应用 Dropout 正则化(regularization),以防止过拟合(overfitting)。 |
Embedding(input_dim, output_dim, ...) | 一个查找表(lookup table),将整数索引(例如词 ID)映射到密集向量嵌入(dense vector embedding)。 |
Flatten() | 将输入展平为 2D 数组,通过折叠除第一个维度(批量维度,batch dimension)外的所有维度来实现。 |
Pool*D (e.g., MaxPool2D, AvgPool2D) | 池化层(Pooling layer,最大池化 max pooling, 平均池化 average pooling),用于特征图(feature map)的下采样(downsampling)。 |
GlobalPool*D (e.g., GlobalAvgPool2D) | 全局池化层(Global pooling layer),将每个特征图缩减为一个单一值。 |
实现示例:Block 和 HybridBlock
Section titled “实现示例:Block 和 HybridBlock”使用 gluon.nn.Block 定义自定义模型:
import mxnet as mxfrom mxnet import nd, gluonfrom mxnet.gluon import nn
class SimpleModel(nn.Block): def __init__(self, **kwargs): super(SimpleModel, self).__init__(**kwargs) with self.name_scope(): # Organizes parameter names self.dense0 = nn.Dense(20) self.dense1 = nn.Dense(20)
def forward(self, x): x = nd.relu(self.dense0(x)) return nd.relu(self.dense1(x))
model_block = SimpleModel()model_block.initialize(ctx=mx.cpu()) # Initialize parametersoutput_block = model_block(nd.zeros((5, 10), ctx=mx.cpu())) # Pass dummy data
print("Output from SimpleModel (Block):")print(output_block.asnumpy())输出(由于 ReLU 和首次传递时可能采用零初始化或小值初始化,所有值均为零):
Output from SimpleModel (Block):[[0. 0. ... 0. 0.] [0. 0. ... 0. 0.] ... [0. 0. ... 0. 0.] [0. 0. ... 0. 0.]] (Shape: 5x20)使用 gluon.nn.HybridBlock 定义自定义模型,以便进行潜在的符号式优化:
class HybridSimpleModel(nn.HybridBlock): def __init__(self, **kwargs): super(HybridSimpleModel, self).__init__(**kwargs) with self.name_scope(): self.dense0 = nn.Dense(30) self.dense1 = nn.Dense(30)
def hybrid_forward(self, F, x): # F is mx.nd or mx.sym depending on whether hybridized x = F.relu(self.dense0(x)) return F.relu(self.dense1(x))
model_hybrid = HybridSimpleModel()model_hybrid.initialize(ctx=mx.cpu())
# Run imperatively firstoutput_hybrid_imperative = model_hybrid(nd.zeros((5, 15), ctx=mx.cpu()))print("\nOutput from HybridSimpleModel (imperative):")print(output_hybrid_imperative.asnumpy())
# Hybridize the model (convert to symbolic graph for subsequent calls)model_hybrid.hybridize()output_hybrid_symbolic = model_hybrid(nd.zeros((5, 15), ctx=mx.cpu()))print("\nOutput from HybridSimpleModel (hybridized/symbolic):")print(output_hybrid_symbolic.asnumpy())输出(形状为 5x30,值可能为零):
Output from HybridSimpleModel (imperative):[[0. 0. ... 0. 0.] ... [0. 0. ... 0. 0.]]
Output from HybridSimpleModel (hybridized/symbolic):[[0. 0. ... 0. 0.] ... [0. 0. ... 0. 0.]]gluon.rnn:循环神经网络层
Section titled “gluon.rnn:循环神经网络层”mxnet.gluon.rnn 模块提供了用于构建循环神经网络(Recurrent Neural Network,RNN)的层,适用于序列数据处理。
关键 RNN 层和单元(Cell)
Section titled “关键 RNN 层和单元(Cell)”| 类名 (Class Name) | 描述 (Description) |
|---|---|
RNNCell (Base class) | RNN 单元的抽象基类。 |
RNN(hidden_size, num_layers, activation='relu', ...) | 多层 Elman RNN。activation 可以是 ‘tanh’ 或 ‘relu’。 |
LSTMCell(hidden_size, ...) | 长短期记忆(Long Short-Term Memory,LSTM)单元。 |
LSTM(hidden_size, num_layers, ...) | 多层 LSTM 网络。 |
GRUCell(hidden_size, ...) | 门控循环单元(Gated Recurrent Unit,GRU)单元。 |
GRU(hidden_size, num_layers, ...) | 多层 GRU 网络。 |
BidirectionalCell(l_cell, r_cell) | 封装两个单元以创建一个双向循环层(Bidirectional recurrent layer)。 |
DropoutCell(rate, ...) | 将 Dropout 应用于循环单元的输入或输出。 |
实现示例:GRU 和 LSTM
Section titled “实现示例:GRU 和 LSTM”使用 gluon.rnn.GRU:
gru_layer = gluon.rnn.GRU(hidden_size=100, num_layers=2)gru_layer.initialize()
# Input sequence: (sequence_length, batch_size, input_feature_dim)input_seq_gru = nd.random.uniform(shape=(5, 3, 10))
# Forward pass without initial state (defaults to zeros)output_seq_gru, last_states_gru = gru_layer(input_seq_gru)print("GRU Output sequence shape:", output_seq_gru.shape)print("GRU Last hidden states shape:", last_states_gru[0].shape) # GRU state is a list containing one tensor输出:
GRU Output sequence shape: (5, 3, 100)GRU Last hidden states shape: (2, 3, 100) (num_layers, batch_size, hidden_size)使用 gluon.rnn.LSTM:
lstm_layer = gluon.rnn.LSTM(hidden_size=120, num_layers=1)lstm_layer.initialize()
input_seq_lstm = nd.random.uniform(shape=(8, 4, 20))
# LSTM requires a list of initial states: [initial_hidden_state, initial_cell_state]# Shapes: (num_layers, batch_size, hidden_size)initial_h_lstm = nd.zeros((1, 4, 120))initial_c_lstm = nd.zeros((1, 4, 120))initial_states_lstm = [initial_h_lstm, initial_c_lstm]
output_seq_lstm, last_states_lstm_tuple = lstm_layer(input_seq_lstm, initial_states_lstm)print("\nLSTM Output sequence shape:", output_seq_lstm.shape)print("LSTM Last hidden state shape:", last_states_lstm_tuple[0].shape)print("LSTM Last cell state shape:", last_states_lstm_tuple[1].shape)输出:
LSTM Output sequence shape: (8, 4, 120)LSTM Last hidden state shape: (1, 4, 120)LSTM Last cell state shape: (1, 4, 120)Gluon 中训练过程中必不可少的模块:
gluon.loss:损失函数
Section titled “gluon.loss:损失函数”mxnet.gluon.loss 模块(通常别名为 gloss)提供了各种预定义的损失函数(loss function),用于评估模型性能并计算训练所需的梯度(gradient)。
常用损失函数
Section titled “常用损失函数”| 类名 (Class Name) | 描述 (Description) |
|---|---|
Loss (Base class) | 所有损失函数的基类。 |
L2Loss() | 计算均方误差(Mean Squared Error,MSE):0.5 * sum((pred - label)^2) / N。0.5 系数在求导时为了方便很常见。 |
L1Loss() | 计算平均绝对误差(Mean Absolute Error,MAE):sum(|pred - label|) / N。 |
SigmoidBinaryCrossEntropyLoss() | 计算二分类(binary classification)问题的交叉熵损失(cross-entropy loss)。如果 from_sigmoid=False(默认值),则先对预测值应用 sigmoid 函数。 |
SoftmaxCrossEntropyLoss() | 计算多分类(multi-class classification)的交叉熵损失。默认情况下对预测值应用 softmax 函数。 |
KLDivLoss() | 计算 Kullback-Leibler 散度损失(KLDiv loss)。 |
CTCLoss() | 连接主义时间分类(Connectionist Temporal Classification,CTC)损失,用于序列到序列(sequence-to-sequence)任务,例如语音识别。 |
HuberLoss() | 一种鲁棒(robust)的损失函数,对于小误差是二次方,对于大误差是线性(比 L2Loss 对异常值 less sensitive)。 |
例如,L2Loss(均方误差)计算的是预测值(pred)与真实标签(label)之间平方差的平均值。一个常见的公式是 L = (1/N) * sum_i (pred_i - label_i)^2,其中 N 是样本或元素的数量。Gluon 的 L2Loss 默认可能使用 0.5 * sum(...) / batch_size。
gluon.Parameter:模型参数
Section titled “gluon.Parameter:模型参数”mxnet.gluon.Parameter 是一个容器类,用于保存 Block 的可学习参数(权重 weights 和偏置 biases)。它管理这些参数的数据、梯度、初始化以及上下文(context,即 CPU/GPU)。
关键 Parameter 方法
Section titled “关键 Parameter 方法”| 方法 (Method) | 描述 (Description) |
|---|---|
data(ctx=None) | 在指定的上下文(或默认上下文)上返回参数的数据 NDArray。 |
grad(ctx=None) | 在指定的上下文上返回参数的梯度 NDArray。 |
initialize(init=None, ctx=None, default_init=init.Uniform(), force_reinit=False) | 在指定的上下文(一个或多个)上初始化参数的数据和梯度数组。 |
list_data() | 返回数据 NDArray 列表,每个列表项对应参数初始化所在的每个上下文。 |
list_grad() | 返回梯度 NDArray 列表,每个列表项对应参数初始化所在的每个上下文。 |
zero_grad() | 将所有上下文上的梯度缓冲区设置为零。如果在典型的训练循环中累积梯度,则在调用 backward() 之前调用此方法。 |
示例:创建并初始化一个参数。
weight_param = gluon.Parameter('my_weight', shape=(2, 2), init=mx.init.Normal(0.1))weight_param.initialize(ctx=mx.cpu(0))
print("Initialized weight_param data:")print(weight_param.data().asnumpy())print("\nInitialized weight_param grad (should be zeros):")print(weight_param.grad().asnumpy())输出(数据值随机,梯度为零):
Initialized weight_param data:[[-0.09423235 0.04523997] [-0.00682008 0.10888068]]
Initialized weight_param grad (should be zeros):[[0. 0.] [0. 0.]]gluon.Trainer:优化参数
Section titled “gluon.Trainer:优化参数”mxnet.gluon.Trainer 类简化了使用优化器(optimizer,例如 SGD、Adam)和计算出的梯度来更新模型参数的过程。
关键 Trainer 方法
Section titled “关键 Trainer 方法”| 方法 (Method) | 描述 (Description) |
|---|---|
Trainer(params, optimizer, optimizer_params=None, kvstore=None, ...) | 构造函数(Constructor)。params 是参数的集合(例如来自 net.collect_params())。optimizer 是优化器的名称(例如 ‘sgd’、‘adam’)或 mx.optimizer.Optimizer 实例。 |
step(batch_size, ignore_stale_grad=False) | 执行一步参数更新。应在 loss.backward() 之后且在 autograd.record() 作用域之外调用此方法。在应用优化器更新之前,它会将梯度按 1.0/batch_size 进行缩放。 |
set_learning_rate(lr) | 更新优化器的学习率(learning rate)。 |
save_states(fname) / load_states(fname) | 保存/加载 Trainer 的状态(例如优化器动量,optimizer moments),对于恢复训练很有用。 |
典型的训练循环片段:
# net = SimpleModel() or HybridSimpleModel()# net.initialize()# loss_fn = gluon.loss.L2Loss()# trainer = gluon.Trainer(net.collect_params(), 'sgd', {'learning_rate': 0.1})# x_batch, y_batch = ... # get a batch of data# with autograd.record():# predictions = net(x_batch)# loss_value = loss_fn(predictions, y_batch)# loss_value.backward() # Compute gradients# trainer.step(batch_size) # Update parameters数据处理模块
Section titled “数据处理模块”Gluon 提供了用于高效加载和管理数据集的实用工具。
gluon.data:数据集(Dataset)和数据加载器(DataLoader)
Section titled “gluon.data:数据集(Dataset)和数据加载器(DataLoader)”mxnet.gluon.data 模块提供了用于创建、转换和以 mini-batch 形式加载数据的工具。
| 类名 (Class Name) | 描述 (Description) |
|---|---|
Dataset (Abstract class) | 数据集的基类。自定义数据集应继承此类。 |
ArrayDataset(*args) | 从一个或多个 NDArray 或 NumPy 数组创建的数据集。每个参数被视为样本的一个组成部分(例如数据、标签)。 |
Sampler (Base class) | 从数据集中采样索引的策略基类。 |
SequentialSampler(length) | 按顺序从 0 到 length-1 采样元素。 |
RandomSampler(length) | 随机、无放回地采样元素。 |
BatchSampler(sampler, batch_size, last_batch='keep') | 包装另一个采样器以生成 mini-batch 的索引。 |
DataLoader(dataset, batch_size=None, shuffle=False, sampler=None, batch_sampler=None, num_workers=0, ...) | 以 mini-batch 形式从 Dataset 加载数据。处理数据混洗(shuffling)、使用 num_workers 进行并行加载。 |
示例:使用 ArrayDataset 和 DataLoader:
num_samples = 100features = nd.random.normal(shape=(num_samples, 10))labels = nd.random.uniform(shape=(num_samples, 1))
# Create a dataset from features and labelsdataset = gluon.data.ArrayDataset(features, labels)
# Create a DataLoader to iterate over mini-batchesbatch_size = 10data_loader = gluon.data.DataLoader(dataset, batch_size=batch_size, shuffle=True)
# Iterate over the data_loader in a training loop# for i, (data_batch, label_batch) in enumerate(data_loader):# print(f"Batch {i}: data shape {data_batch.shape}, label shape {label_batch.shape}")# if i == 1: break # Show first two batchesprint(f"DataLoader created for {len(dataset)} samples with batch_size {batch_size}.")gluon.data.vision.datasets:视觉数据集
Section titled “gluon.data.vision.datasets:视觉数据集”此模块提供了访问常见计算机视觉(computer vision)数据集的便捷方式。
预定义视觉数据集
Section titled “预定义视觉数据集”| 类名 (Class Name) | 描述 (Description) |
|---|---|
MNIST([root, train, transform]) | MNIST 手写数字数据集。 |
FashionMNIST([root, train, transform]) | Fashion-MNIST 数据集(可直接替代 MNIST)。 |
CIFAR10([root, train, transform]) | CIFAR-10 图像分类数据集(10 个类别)。 |
CIFAR100([root, fine_label, train, transform]) | CIFAR-100 图像分类数据集(100 个类别)。 |
ImageRecordDataset(filename, ...) | 从存储在 MXNet RecordIO(.rec)格式文件中的图像创建的数据集。 |
ImageFolderDataset(root, ...) | 从存储在文件夹结构(例如 root/class_name/image.jpg)中的图像创建的数据集。 |
ImageListDataset(root, imglistfile, ...) | 从列表文件(.lst)中指定的图像创建的数据集。 |
ImageListDataset 文件格式(.lst)示例:每行通常是 index label path_to_image
# Example content of an image_list.lst file:# 0<tab>0<tab>path/to/cat_image_001.jpg# 1<tab>0<tab>path/to/cat_image_002.jpg# 2<tab>1<tab>path/to/dog_image_001.jpg# ... (index, label, image_path)实用工具模块
Section titled “实用工具模块”Gluon 还包含用于各种任务的实用工具函数。
gluon.utils:训练实用工具
Section titled “gluon.utils:训练实用工具”mxnet.gluon.utils 模块包含辅助函数,特别是用于数据并行(data parallelism)和文件处理。
关键实用工具函数
Section titled “关键实用工具函数”| 函数 (Function) | 描述 (Description) |
|---|---|
split_data(data, num_slice, batch_axis=0, ...) | 沿 batch_axis 将一个 NDArray(或 NDArray 列表)分割成 num_slice 片。不将数据移动到上下文。 |
split_and_load(data, ctx_list, batch_axis=0, ...) | 类似于 split_data 分割数据,然后将每片数据加载到 ctx_list 中对应的上下文上。对于单机多 GPU 训练至关重要。 |
clip_global_norm(arrays, max_norm, ...) | 重新缩放 NDArray 列表(通常是梯度),使其总 L2 范数小于或等于 max_norm。有助于防止梯度爆炸(exploding gradients)。 |
download(url, path=None, overwrite=False, sha1_hash=None, ...) | 从 URL 下载文件。可以使用 sha1_hash 验证文件的完整性(integrity)。 |