Skip to content

Apache MXNet - Python API gluon

MXNet Gluon API 提供了一个用户友好的、命令式(imperative)接口,用于设计、训练和部署深度学习模型。它通过其混合(hybrid)功能,将动态框架的易用性与符号式执行(symbolic execution)的性能结合起来。

本章将探索 mxnet.gluon 包中的关键模块。

mxnet.gluon.nn 模块是构建模型的核心,提供了丰富的预定义神经网络层(Layer)集合。

gluon.nn 中的一些重要类(class):

类名 (Class Name)描述 (Description)
Block([prefix, params])Gluon 中所有神经网络层和模型的基类。通过继承 Block 可以构建自定义模型。
HybridBlock([prefix, params])Block 的子类,支持命令式执行和符号图编译(混合化,hybridization),用于性能优化。
Sequential([prefix, params])用于按顺序堆叠层(Layer)的容器。HybridSequential 是其对应的混合版本。
Dense(units, activation=None, use_bias=True, ...)标准的ᵉ全连接层(fully connected layer,或称密集层,dense layer)。
Conv1D, Conv2D, Conv3D分别对应 1D、2D 和 3D 卷积层(convolution layer)。参数包括 channels(通道数)、kernel_size(卷积核大小)、strides(步长)、padding(填充)、activation(激活函数)。
Conv*DTranspose转置卷积层(Transposed convolution layer,常被称为反卷积层,deconvolution layer),用于上采样(upsampling)。
Activation(activation_type)应用指定的激活函数(activation function),例如 ‘relu’、‘sigmoid’、‘tanh’。
BatchNorm([axis, momentum, epsilon, ...])批量归一化层(Batch normalization layer),用于稳定训练并提高泛化能力。
Dropout(rate, axes=())应用 Dropout 正则化(regularization),以防止过拟合(overfitting)。
Embedding(input_dim, output_dim, ...)一个查找表(lookup table),将整数索引(例如词 ID)映射到密集向量嵌入(dense vector embedding)。
Flatten()将输入展平为 2D 数组,通过折叠除第一个维度(批量维度,batch dimension)外的所有维度来实现。
Pool*D (e.g., MaxPool2D, AvgPool2D)池化层(Pooling layer,最大池化 max pooling, 平均池化 average pooling),用于特征图(feature map)的下采样(downsampling)。
GlobalPool*D (e.g., GlobalAvgPool2D)全局池化层(Global pooling layer),将每个特征图缩减为一个单一值。

使用 gluon.nn.Block 定义自定义模型:

import mxnet as mx
from mxnet import nd, gluon
from mxnet.gluon import nn
class SimpleModel(nn.Block):
def __init__(self, **kwargs):
super(SimpleModel, self).__init__(**kwargs)
with self.name_scope(): # Organizes parameter names
self.dense0 = nn.Dense(20)
self.dense1 = nn.Dense(20)
def forward(self, x):
x = nd.relu(self.dense0(x))
return nd.relu(self.dense1(x))
model_block = SimpleModel()
model_block.initialize(ctx=mx.cpu()) # Initialize parameters
output_block = model_block(nd.zeros((5, 10), ctx=mx.cpu())) # Pass dummy data
print("Output from SimpleModel (Block):")
print(output_block.asnumpy())

输出(由于 ReLU 和首次传递时可能采用零初始化或小值初始化,所有值均为零):

Output from SimpleModel (Block):
[[0. 0. ... 0. 0.]
[0. 0. ... 0. 0.]
...
[0. 0. ... 0. 0.]
[0. 0. ... 0. 0.]] (Shape: 5x20)

使用 gluon.nn.HybridBlock 定义自定义模型,以便进行潜在的符号式优化:

class HybridSimpleModel(nn.HybridBlock):
def __init__(self, **kwargs):
super(HybridSimpleModel, self).__init__(**kwargs)
with self.name_scope():
self.dense0 = nn.Dense(30)
self.dense1 = nn.Dense(30)
def hybrid_forward(self, F, x):
# F is mx.nd or mx.sym depending on whether hybridized
x = F.relu(self.dense0(x))
return F.relu(self.dense1(x))
model_hybrid = HybridSimpleModel()
model_hybrid.initialize(ctx=mx.cpu())
# Run imperatively first
output_hybrid_imperative = model_hybrid(nd.zeros((5, 15), ctx=mx.cpu()))
print("\nOutput from HybridSimpleModel (imperative):")
print(output_hybrid_imperative.asnumpy())
# Hybridize the model (convert to symbolic graph for subsequent calls)
model_hybrid.hybridize()
output_hybrid_symbolic = model_hybrid(nd.zeros((5, 15), ctx=mx.cpu()))
print("\nOutput from HybridSimpleModel (hybridized/symbolic):")
print(output_hybrid_symbolic.asnumpy())

输出(形状为 5x30,值可能为零):

Output from HybridSimpleModel (imperative):
[[0. 0. ... 0. 0.]
...
[0. 0. ... 0. 0.]]
Output from HybridSimpleModel (hybridized/symbolic):
[[0. 0. ... 0. 0.]
...
[0. 0. ... 0. 0.]]

mxnet.gluon.rnn 模块提供了用于构建循环神经网络(Recurrent Neural Network,RNN)的层,适用于序列数据处理。

类名 (Class Name)描述 (Description)
RNNCell (Base class)RNN 单元的抽象基类。
RNN(hidden_size, num_layers, activation='relu', ...)多层 Elman RNN。activation 可以是 ‘tanh’ 或 ‘relu’。
LSTMCell(hidden_size, ...)长短期记忆(Long Short-Term Memory,LSTM)单元。
LSTM(hidden_size, num_layers, ...)多层 LSTM 网络。
GRUCell(hidden_size, ...)门控循环单元(Gated Recurrent Unit,GRU)单元。
GRU(hidden_size, num_layers, ...)多层 GRU 网络。
BidirectionalCell(l_cell, r_cell)封装两个单元以创建一个双向循环层(Bidirectional recurrent layer)。
DropoutCell(rate, ...)将 Dropout 应用于循环单元的输入或输出。

使用 gluon.rnn.GRU:

gru_layer = gluon.rnn.GRU(hidden_size=100, num_layers=2)
gru_layer.initialize()
# Input sequence: (sequence_length, batch_size, input_feature_dim)
input_seq_gru = nd.random.uniform(shape=(5, 3, 10))
# Forward pass without initial state (defaults to zeros)
output_seq_gru, last_states_gru = gru_layer(input_seq_gru)
print("GRU Output sequence shape:", output_seq_gru.shape)
print("GRU Last hidden states shape:", last_states_gru[0].shape) # GRU state is a list containing one tensor

输出:

GRU Output sequence shape: (5, 3, 100)
GRU Last hidden states shape: (2, 3, 100) (num_layers, batch_size, hidden_size)

使用 gluon.rnn.LSTM:

lstm_layer = gluon.rnn.LSTM(hidden_size=120, num_layers=1)
lstm_layer.initialize()
input_seq_lstm = nd.random.uniform(shape=(8, 4, 20))
# LSTM requires a list of initial states: [initial_hidden_state, initial_cell_state]
# Shapes: (num_layers, batch_size, hidden_size)
initial_h_lstm = nd.zeros((1, 4, 120))
initial_c_lstm = nd.zeros((1, 4, 120))
initial_states_lstm = [initial_h_lstm, initial_c_lstm]
output_seq_lstm, last_states_lstm_tuple = lstm_layer(input_seq_lstm, initial_states_lstm)
print("\nLSTM Output sequence shape:", output_seq_lstm.shape)
print("LSTM Last hidden state shape:", last_states_lstm_tuple[0].shape)
print("LSTM Last cell state shape:", last_states_lstm_tuple[1].shape)

输出:

LSTM Output sequence shape: (8, 4, 120)
LSTM Last hidden state shape: (1, 4, 120)
LSTM Last cell state shape: (1, 4, 120)

Gluon 中训练过程中必不可少的模块:

mxnet.gluon.loss 模块(通常别名为 gloss)提供了各种预定义的损失函数(loss function),用于评估模型性能并计算训练所需的梯度(gradient)。

类名 (Class Name)描述 (Description)
Loss (Base class)所有损失函数的基类。
L2Loss()计算均方误差(Mean Squared Error,MSE):0.5 * sum((pred - label)^2) / N。0.5 系数在求导时为了方便很常见。
L1Loss()计算平均绝对误差(Mean Absolute Error,MAE):sum(|pred - label|) / N。
SigmoidBinaryCrossEntropyLoss()计算二分类(binary classification)问题的交叉熵损失(cross-entropy loss)。如果 from_sigmoid=False(默认值),则先对预测值应用 sigmoid 函数。
SoftmaxCrossEntropyLoss()计算多分类(multi-class classification)的交叉熵损失。默认情况下对预测值应用 softmax 函数。
KLDivLoss()计算 Kullback-Leibler 散度损失(KLDiv loss)。
CTCLoss()连接主义时间分类(Connectionist Temporal Classification,CTC)损失,用于序列到序列(sequence-to-sequence)任务,例如语音识别。
HuberLoss()一种鲁棒(robust)的损失函数,对于小误差是二次方,对于大误差是线性(比 L2Loss 对异常值 less sensitive)。

例如,L2Loss(均方误差)计算的是预测值(pred)与真实标签(label)之间平方差的平均值。一个常见的公式是 L = (1/N) * sum_i (pred_i - label_i)^2,其中 N 是样本或元素的数量。Gluon 的 L2Loss 默认可能使用 0.5 * sum(...) / batch_size。

mxnet.gluon.Parameter 是一个容器类,用于保存 Block 的可学习参数(权重 weights 和偏置 biases)。它管理这些参数的数据、梯度、初始化以及上下文(context,即 CPU/GPU)。

方法 (Method)描述 (Description)
data(ctx=None)在指定的上下文(或默认上下文)上返回参数的数据 NDArray。
grad(ctx=None)在指定的上下文上返回参数的梯度 NDArray。
initialize(init=None, ctx=None, default_init=init.Uniform(), force_reinit=False)在指定的上下文(一个或多个)上初始化参数的数据和梯度数组。
list_data()返回数据 NDArray 列表,每个列表项对应参数初始化所在的每个上下文。
list_grad()返回梯度 NDArray 列表,每个列表项对应参数初始化所在的每个上下文。
zero_grad()将所有上下文上的梯度缓冲区设置为零。如果在典型的训练循环中累积梯度,则在调用 backward() 之前调用此方法。

示例:创建并初始化一个参数。

weight_param = gluon.Parameter('my_weight', shape=(2, 2), init=mx.init.Normal(0.1))
weight_param.initialize(ctx=mx.cpu(0))
print("Initialized weight_param data:")
print(weight_param.data().asnumpy())
print("\nInitialized weight_param grad (should be zeros):")
print(weight_param.grad().asnumpy())

输出(数据值随机,梯度为零):

Initialized weight_param data:
[[-0.09423235 0.04523997]
[-0.00682008 0.10888068]]
Initialized weight_param grad (should be zeros):
[[0. 0.]
[0. 0.]]

mxnet.gluon.Trainer 类简化了使用优化器(optimizer,例如 SGD、Adam)和计算出的梯度来更新模型参数的过程。

方法 (Method)描述 (Description)
Trainer(params, optimizer, optimizer_params=None, kvstore=None, ...)构造函数(Constructor)。params 是参数的集合(例如来自 net.collect_params())。optimizer 是优化器的名称(例如 ‘sgd’、‘adam’)或 mx.optimizer.Optimizer 实例。
step(batch_size, ignore_stale_grad=False)执行一步参数更新。应在 loss.backward() 之后且在 autograd.record() 作用域之外调用此方法。在应用优化器更新之前,它会将梯度按 1.0/batch_size 进行缩放。
set_learning_rate(lr)更新优化器的学习率(learning rate)。
save_states(fname) / load_states(fname)保存/加载 Trainer 的状态(例如优化器动量,optimizer moments),对于恢复训练很有用。

典型的训练循环片段:

# net = SimpleModel() or HybridSimpleModel()
# net.initialize()
# loss_fn = gluon.loss.L2Loss()
# trainer = gluon.Trainer(net.collect_params(), 'sgd', {'learning_rate': 0.1})
# x_batch, y_batch = ... # get a batch of data
# with autograd.record():
# predictions = net(x_batch)
# loss_value = loss_fn(predictions, y_batch)
# loss_value.backward() # Compute gradients
# trainer.step(batch_size) # Update parameters

Gluon 提供了用于高效加载和管理数据集的实用工具。

gluon.data:数据集(Dataset)和数据加载器(DataLoader)

Section titled “gluon.data:数据集(Dataset)和数据加载器(DataLoader)”

mxnet.gluon.data 模块提供了用于创建、转换和以 mini-batch 形式加载数据的工具。

类名 (Class Name)描述 (Description)
Dataset (Abstract class)数据集的基类。自定义数据集应继承此类。
ArrayDataset(*args)从一个或多个 NDArray 或 NumPy 数组创建的数据集。每个参数被视为样本的一个组成部分(例如数据、标签)。
Sampler (Base class)从数据集中采样索引的策略基类。
SequentialSampler(length)按顺序从 0 到 length-1 采样元素。
RandomSampler(length)随机、无放回地采样元素。
BatchSampler(sampler, batch_size, last_batch='keep')包装另一个采样器以生成 mini-batch 的索引。
DataLoader(dataset, batch_size=None, shuffle=False, sampler=None, batch_sampler=None, num_workers=0, ...)以 mini-batch 形式从 Dataset 加载数据。处理数据混洗(shuffling)、使用 num_workers 进行并行加载。

示例:使用 ArrayDataset 和 DataLoader:

num_samples = 100
features = nd.random.normal(shape=(num_samples, 10))
labels = nd.random.uniform(shape=(num_samples, 1))
# Create a dataset from features and labels
dataset = gluon.data.ArrayDataset(features, labels)
# Create a DataLoader to iterate over mini-batches
batch_size = 10
data_loader = gluon.data.DataLoader(dataset, batch_size=batch_size, shuffle=True)
# Iterate over the data_loader in a training loop
# for i, (data_batch, label_batch) in enumerate(data_loader):
# print(f"Batch {i}: data shape {data_batch.shape}, label shape {label_batch.shape}")
# if i == 1: break # Show first two batches
print(f"DataLoader created for {len(dataset)} samples with batch_size {batch_size}.")

gluon.data.vision.datasets:视觉数据集

Section titled “gluon.data.vision.datasets:视觉数据集”

此模块提供了访问常见计算机视觉(computer vision)数据集的便捷方式。

类名 (Class Name)描述 (Description)
MNIST([root, train, transform])MNIST 手写数字数据集。
FashionMNIST([root, train, transform])Fashion-MNIST 数据集(可直接替代 MNIST)。
CIFAR10([root, train, transform])CIFAR-10 图像分类数据集(10 个类别)。
CIFAR100([root, fine_label, train, transform])CIFAR-100 图像分类数据集(100 个类别)。
ImageRecordDataset(filename, ...)从存储在 MXNet RecordIO(.rec)格式文件中的图像创建的数据集。
ImageFolderDataset(root, ...)从存储在文件夹结构(例如 root/class_name/image.jpg)中的图像创建的数据集。
ImageListDataset(root, imglistfile, ...)从列表文件(.lst)中指定的图像创建的数据集。

ImageListDataset 文件格式(.lst)示例:每行通常是 index label path_to_image

# Example content of an image_list.lst file:
# 0<tab>0<tab>path/to/cat_image_001.jpg
# 1<tab>0<tab>path/to/cat_image_002.jpg
# 2<tab>1<tab>path/to/dog_image_001.jpg
# ... (index, label, image_path)

Gluon 还包含用于各种任务的实用工具函数。

mxnet.gluon.utils 模块包含辅助函数,特别是用于数据并行(data parallelism)和文件处理。

函数 (Function)描述 (Description)
split_data(data, num_slice, batch_axis=0, ...)沿 batch_axis 将一个 NDArray(或 NDArray 列表)分割成 num_slice 片。不将数据移动到上下文。
split_and_load(data, ctx_list, batch_axis=0, ...)类似于 split_data 分割数据,然后将每片数据加载到 ctx_list 中对应的上下文上。对于单机多 GPU 训练至关重要。
clip_global_norm(arrays, max_norm, ...)重新缩放 NDArray 列表(通常是梯度),使其总 L2 范数小于或等于 max_norm。有助于防止梯度爆炸(exploding gradients)。
download(url, path=None, overwrite=False, sha1_hash=None, ...)从 URL 下载文件。可以使用 sha1_hash 验证文件的完整性(integrity)。