Skip to content

Apache MXNet - Gluon

Gluon API (mxnet.gluon) 是 Apache MXNet 中的一个高级接口,旨在提供易用性、灵活性以及深度学习模型的快速原型开发。它提供了命令式编程风格,类似于标准的 Python,使开发者能够直观地进行开发。虽然易于使用,Gluon 通过允许模型被 hybridized(即转换为符号图进行优化)也可以实现高性能。

gluon.Block 是在 Gluon 中创建神经网络层和模型的基本构建单元。任何神经网络层或构成更大网络组件的一组层,通常都是 Block 的子类。

Block 的主要特点:

  • 它可以包含其他的 Block(允许分层的模型结构)。
  • 它持有与其定义的层相关的参数(weights 和 biases)。
  • 它定义了一个 forward 方法(对于 HybridBlock 则是 hybrid_forward 方法),该方法指定了 Block 执行的计算。

示例:使用 nn.Sequential 构建一个简单的多层感知机 (MLP)

Section titled “示例:使用 nn.Sequential 构建一个简单的多层感知机 (MLP)”

gluon.nn.Sequential 是一个特殊的 Block,它按顺序堆叠其他的 Block,其中一个 Block 的输出作为下一个 Block 的输入。

import mxnet as mx
from mxnet import nd, gluon, autograd
from mxnet.gluon import nn # nn is a submodule of gluon containing predefined layers
mx.random.seed(0) # for reproducibility
# Dummy input data (batch_size=2, num_features=20)
x = nd.random.uniform(shape=(2, 20), ctx=mx.cpu())
# Define a sequential MLP
N_net = nn.Sequential()
with N_net.name_scope(): # Organizes parameter names
N_net.add(nn.Dense(256, activation='relu')) # Fully connected layer with 256 units and ReLU activation
N_net.add(nn.Dense(10)) # Output layer with 10 units (e.g., for 10-class classification)
# Initialize parameters (weights and biases)
# This needs to be done before the first forward pass if not using deferred initialization with actual data.
# Alternatively, with deferred initialization (default for many layers), shapes are inferred on first forward pass.
N_net.initialize(ctx=mx.cpu()) # Initialize on CPU context
# Perform a forward pass
output = N_net(x)
print(output)

输出(实际值取决于随机初始化):

<html>
<body>
<p>
[[ 0.07308105 0.00406009 -0.00294911 0.0370597 0.00400224 0.008209
-0.02912429 -0.05099454 -0.0040519 0.01902002]
[ 0.09712769 0.00383658 0.01095401 0.02602063 0.02011784 0.00934159
-0.01897359 -0.01587484 -0.01876913 0.02037334]]
<NDArray 2x10 @cpu(0)>
</p>
</body>
</html>

使用 Block(如 nn.Sequential 或自定义 Block)所涉及的步骤:

  1. 定义 Block:实例化它或定义其架构(例如,通过向 nn.Sequential 添加层)。
  2. 初始化参数:调用 .initialize() 来设置 weights 和 biases。MXNet 支持各种 initializer(例如,Xavier, Normal)。如果初始化时未指定,shape inference 通常会在第一次 forward pass 时自动发生。
  3. Forward Pass:将输入数据 (NDArray) 作为函数调用(例如,N_net(x))通过 Block。这将执行定义的计算。
  4. Backward Pass(训练期间):使用 autograd.record() 记录计算,然后对 loss 的标量值调用 .backward() 来计算 Block 参数的 gradient。

我们可以实现一个简化版的 Sequential 来理解其核心逻辑:

class MySequential(nn.Block):
def __init__(self, **kwargs):
super(MySequential, self).__init__(**kwargs)
# self._children is an OrderedDict in the actual nn.Block to store sub-blocks
def add(self, block):
# In a real Block, self._children stores blocks with unique names.
# For simplicity, we use a list here. Gluon uses an OrderedDict.
if not hasattr(self, '_added_blocks'):
self._added_blocks = []
self._added_blocks.append(block)
self.register_child(block) # Important for parameter collection etc.
def forward(self, x):
for block in self._added_blocks: # Iterate through added blocks
x = block(x)
return x
# Using MySequential
x_custom = nd.random.uniform(shape=(2, 20), ctx=mx.cpu())
N_net_custom = MySequential()
with N_net_custom.name_scope():
N_net_custom.add(nn.Dense(256, activation='relu'))
N_net_custom.add(nn.Dense(10))
N_net_custom.initialize(ctx=mx.cpu())
output_custom = N_net_custom(x_custom)
print(f"Output from MySequential (should be similar shape and type):\n{output_custom}")

输出(实际值因单独初始化而不同):

<html>
<body>
<p>
Output from MySequential (should be similar shape and type):
[[ 0.01903114 0.02831168 0.01545286 -0.01109749 0.0350692 -0.02966933
-0.01100099 -0.02197874 -0.01193464 0.00189657]
[ 0.0183696 0.02339937 0.01078339 -0.0076901 0.03513341 -0.02404617
-0.00755438 -0.01889096 -0.00740707 0.00199449]]
<NDArray 2x10 @cpu(0)>
</p>
</body>
</html>

对于层不简单地按顺序排列的更复杂架构(例如,带有 skip connection 的网络如 ResNet,或多个输入/输出分支),您需要通过继承 gluon.Block(或 gluon.HybridBlock)来创建自定义 Block。

在自定义 Block 中,您通常需要:

  1. 在 __init__ 方法中将层定义为属性。
  2. 实现 forward(self, x, ...) 方法来定义计算逻辑,使用在 __init__ 中定义的层。
class MLP(nn.Block):
def __init__(self, **kwargs):
super(MLP, self).__init__(**kwargs)
# Define layers within the name_scope to manage parameter names
with self.name_scope():
self.hidden = nn.Dense(256, activation='relu') # Hidden layer
self.output_layer = nn.Dense(10) # Output layer (renamed to avoid conflict with 'output' variable)
def forward(self, x):
# Define the forward computation
hidden_activation = self.hidden(x)
final_output = self.output_layer(hidden_activation)
return final_output
x_mlp = nd.random.uniform(shape=(2, 20), ctx=mx.cpu())
N_net_mlp = MLP()
N_net_mlp.initialize(ctx=mx.cpu())
output_mlp = N_net_mlp(x_mlp)
print(f"Output from custom MLP block:\n{output_mlp}")

输出(值取决于初始化):

<html>
<body>
<p>
Output from custom MLP block:
[[-0.02450891 -0.02139075 -0.04023137 0.01311899 0.00839001 0.01907403
0.00343563 0.00571267 -0.02951591 -0.03287996]
[-0.01569841 -0.01874819 -0.0320272 0.0182126 0.00212092 0.01575342
0.0099024 0.00390883 -0.02429737 -0.02714972]]
<NDArray 2x10 @cpu(0)>
</p>
</body>
</html>

虽然 Gluon 在 gluon.nn 中提供了许多常用层,但您可能需要从头开始实现一个自定义层。这可以通过继承 gluon.Block(用于纯命令式层)或 gluon.HybridBlock(用于可以被 hybridized 的层)来实现。

一个简单的自定义层(不含可学习参数)

Section titled “一个简单的自定义层(不含可学习参数)”

自定义层需要实现 forward 方法。如果它没有可学习参数,__init__ 可能只需要调用父类构造函数。

# A layer that normalizes input to the range [0, 1]
class NormalizationLayer(gluon.Block):
def __init__(self, **kwargs):
super(NormalizationLayer, self).__init__(**kwargs)
def forward(self, x):
# Ensure min_val is not equal to max_val to avoid division by zero
min_val = nd.min(x)
max_val = nd.max(x)
if min_val == max_val:
return nd.zeros_like(x)
return (x - min_val) / (max_val - min_val)
x_norm_layer_input = nd.array([[1, 2, 3], [4, 5, 6]], ctx=mx.cpu(), dtype='float32')
norm_layer = NormalizationLayer()
# No .initialize() needed if there are no parameters, but good practice if it might gain params later.
# norm_layer.initialize(ctx=mx.cpu()) # Would be harmless
output_norm_layer = norm_layer(x_norm_layer_input)
print(f"Output from NormalizationLayer:\n{output_norm_layer}")

输出:

<html>
<body>
<p>
Output from NormalizationLayer:
[[0. 0.2 0.4]
[0.6 0.8 1. ]]
<NDArray 2x3 @cpu(0)>
</p>
</body>
</html>

Hybridization:结合命令式和符号式执行

Section titled “Hybridization:结合命令式和符号式执行”

gluon.HybridBlock 是一种特殊的 Block,它可以在两种模式下运行:命令式(像常规 Block 一样)和符号式。定义 HybridBlock 后,您可以调用其 .hybridize() 方法。这将尝试将 hybrid_forward 方法转换为符号图,然后 MXNet 可以对其进行优化以实现更快的执行。

hybridization 的好处:

  • 灵活性:以命令式方式开发和调试。
  • 性能:通过符号图优化进行部署,以提高速度并降低内存使用。
  • 可移植性:导出符号图以便在其他语言绑定或设备上进行推理。

创建 HybridBlock 时,您需要实现 hybrid_forward(self, F, x, ...) 而不是 forward。F 参数至关重要:它引用 mxnet.ndarray(在命令式模式下运行)或 mxnet.symbol(在调用 .hybridize() 后构建符号图)。所有后端操作都必须使用 F(例如,F.relu(x) 而不是 nd.relu(x))。

class NormalizationHybridLayer(gluon.HybridBlock):
def __init__(self, **kwargs):
super(NormalizationHybridLayer, self).__init__(**kwargs)
def hybrid_forward(self, F, x):
min_val = F.min(x)
max_val = F.max(x)
# Add a small epsilon to prevent division by zero if min_val == max_val
# F.broadcast_xxx operations are common for N-D array compatibility
denominator = F.broadcast_sub(max_val, min_val) + 1e-8
return F.broadcast_div(F.broadcast_sub(x, min_val), denominator)
input_hybrid_norm = nd.array([1, 2, 3, 4, 5, 6], ctx=mx.cpu(), dtype='float32')
layer_hybd = NormalizationHybridLayer()
# layer_hybd.initialize() # Not strictly needed without params, but good practice
# Run imperatively first
output_imperative = layer_hybd(input_hybrid_norm)
print(f"Imperative output:\n{output_imperative}")
# Hybridize the layer
layer_hybd.hybridize()
# Run again (now using symbolic graph internally)
output_hybridized = layer_hybd(input_hybrid_norm)
print(f"Hybridized output:\n{output_hybridized}")

输出:

<html>
<body>
<p>
Imperative output:
[0. 0.2 0.4 0.6 0.8 1. ]
<NDArray 6 @cpu(0)>
Hybridized output:
[0. 0.2 0.4 0.6 0.8 1. ]
<NDArray 6 @cpu(0)>
</p>
</body>
</html>

注意:Hybridization 是将模型编译成符号图的一个步骤。执行时使用 CPU 还是 GPU 是独立于此的,取决于输入数据和层参数的 context。

主要区别在于是否可以 hybridization。Block 始终以命令式方式运行。HybridBlock 在调用 .hybridize() 之前可以以命令式方式运行,之后则以符号式(编译后)方式运行。这要求在 hybrid_forward 中使用 F 参数进行所有 MXNet 操作,以确保与 mx.nd 和 mx.sym 后端兼容。

自定义层,特别是 HybridBlock,可以使用像 gluon.nn.Sequential 或 gluon.nn.HybridSequential 这样的容器与标准层结合使用。HybridSequential 本身就是一个 HybridBlock,用于堆叠其他 HybridBlock。

net = gluon.nn.HybridSequential()
with net.name_scope():
net.add(nn.Dense(5, activation='relu')) # Standard dense layer
net.add(NormalizationHybridLayer()) # Our custom hybrid layer
net.add(nn.Dense(1)) # Another standard dense layer
net.initialize(mx.init.Xavier(magnitude=2.24), ctx=mx.cpu())
# Hybridize the entire network
net.hybridize()
# Dummy input
input_data_for_net = nd.random.uniform(low=-10, high=10, shape=(10, 2), ctx=mx.cpu())
output_from_net = net(input_data_for_net)
print(f"Output from hybrid sequential net (shape {output_from_net.shape}):\n{output_from_net[0:3]}") # Print first 3 rows

输出(前 3 行,实际值会有所不同):

<html>
<body>
<p>
Output from hybrid sequential net (shape (10, 1)):
[[-0.40172195]
[-0.40172195]
[-0.40172195]]
<NDArray 3x1 @cpu(0)>
</p>
</body>
</html>

层通常具有可学习参数(weights, biases)或固定的配置参数。在 Gluon 中,参数通过 self.params 管理,它是每个 Block 中可用的 ParameterDict 对象。

  • 可学习参数:使用 self.params.get(...) 定义。它们的 shape 通常可以被 inferred (allow_deferred_init=True)。它们在训练期间会被更新。
  • 固定参数/Buffer:如果需要,也可以存储,有时作为常量或不可微分参数。

示例:带有可学习 Weights 和固定 Scales 的自定义层

Section titled “示例:带有可学习 Weights 和固定 Scales 的自定义层”
class ScaledFullyConnectedLayer(gluon.HybridBlock):
def __init__(self, hidden_units, fixed_scales_data, **kwargs):
super(ScaledFullyConnectedLayer, self).__init__(**kwargs)
self.hidden_units = hidden_units
with self.name_scope():
# Learnable weights for a fully connected operation
# Shape is (output_units, input_units_or_0_for_inference)
self.weights = self.params.get('weights',
shape=(hidden_units, 0), # Input dim inferred later
allow_deferred_init=True,
init=mx.init.Normal(0.01))
# Fixed scales (not trainable)
self.scales = self.params.get('scales',
shape=fixed_scales_data.shape,
init=mx.init.Constant(fixed_scales_data.asnumpy()),
differentiable=False) # Not updated by optimizer
# Bias (learnable)
self.bias = self.params.get('bias',
shape=(hidden_units,),
init=mx.init.Constant(0),
allow_deferred_init=True)
def hybrid_forward(self, F, x, weights, scales, bias):
# Fully connected operation using F for backend compatibility
fc_out = F.FullyConnected(data=x, weight=weights, bias=bias, num_hidden=self.hidden_units, no_bias=False)
# Element-wise multiplication with scales
scaled_out = F.broadcast_mul(fc_out, scales)
return scaled_out
# Example usage
num_features_in = 10
num_hidden_units = 5
fixed_scales = nd.array([0.1, 0.5, 1.0, 0.5, 0.1], ctx=mx.cpu()) # Scales for each hidden unit
scaled_fc_layer = ScaledFullyConnectedLayer(hidden_units=num_hidden_units, fixed_scales_data=fixed_scales)
scaled_fc_layer.initialize(ctx=mx.cpu())
dummy_input_fc = nd.random.normal(shape=(3, num_features_in), ctx=mx.cpu())
output_scaled_fc = scaled_fc_layer(dummy_input_fc)
print(f"Input shape: {dummy_input_fc.shape}")
print(f"Output shape from ScaledFullyConnectedLayer: {output_scaled_fc.shape}")
print(f"Parameters in layer: {scaled_fc_layer.collect_params()}")