Apache MXNet - Gluon
Apache MXNet - Gluon API
Section titled “Apache MXNet - Gluon API”Gluon API (mxnet.gluon) 是 Apache MXNet 中的一个高级接口,旨在提供易用性、灵活性以及深度学习模型的快速原型开发。它提供了命令式编程风格,类似于标准的 Python,使开发者能够直观地进行开发。虽然易于使用,Gluon 通过允许模型被 hybridized(即转换为符号图进行优化)也可以实现高性能。
Block:Gluon 模型的构建单元
Section titled “Block:Gluon 模型的构建单元”gluon.Block 是在 Gluon 中创建神经网络层和模型的基本构建单元。任何神经网络层或构成更大网络组件的一组层,通常都是 Block 的子类。
Block 的主要特点:
- 它可以包含其他的
Block(允许分层的模型结构)。 - 它持有与其定义的层相关的参数(weights 和 biases)。
- 它定义了一个
forward方法(对于HybridBlock则是hybrid_forward方法),该方法指定了 Block 执行的计算。
示例:使用 nn.Sequential 构建一个简单的多层感知机 (MLP)
Section titled “示例:使用 nn.Sequential 构建一个简单的多层感知机 (MLP)”gluon.nn.Sequential 是一个特殊的 Block,它按顺序堆叠其他的 Block,其中一个 Block 的输出作为下一个 Block 的输入。
import mxnet as mxfrom mxnet import nd, gluon, autogradfrom mxnet.gluon import nn # nn is a submodule of gluon containing predefined layers
mx.random.seed(0) # for reproducibility
# Dummy input data (batch_size=2, num_features=20)x = nd.random.uniform(shape=(2, 20), ctx=mx.cpu())
# Define a sequential MLPN_net = nn.Sequential()with N_net.name_scope(): # Organizes parameter names N_net.add(nn.Dense(256, activation='relu')) # Fully connected layer with 256 units and ReLU activation N_net.add(nn.Dense(10)) # Output layer with 10 units (e.g., for 10-class classification)
# Initialize parameters (weights and biases)# This needs to be done before the first forward pass if not using deferred initialization with actual data.# Alternatively, with deferred initialization (default for many layers), shapes are inferred on first forward pass.N_net.initialize(ctx=mx.cpu()) # Initialize on CPU context
# Perform a forward passoutput = N_net(x)print(output)输出(实际值取决于随机初始化):
<html> <body> <p> [[ 0.07308105 0.00406009 -0.00294911 0.0370597 0.00400224 0.008209 -0.02912429 -0.05099454 -0.0040519 0.01902002] [ 0.09712769 0.00383658 0.01095401 0.02602063 0.02011784 0.00934159 -0.01897359 -0.01587484 -0.01876913 0.02037334]] <NDArray 2x10 @cpu(0)> </p> </body></html>使用 Block(如 nn.Sequential 或自定义 Block)所涉及的步骤:
- 定义 Block:实例化它或定义其架构(例如,通过向
nn.Sequential添加层)。 - 初始化参数:调用
.initialize()来设置 weights 和 biases。MXNet 支持各种 initializer(例如,Xavier, Normal)。如果初始化时未指定,shape inference 通常会在第一次 forward pass 时自动发生。 - Forward Pass:将输入数据 (NDArray) 作为函数调用(例如,
N_net(x))通过 Block。这将执行定义的计算。 - Backward Pass(训练期间):使用
autograd.record()记录计算,然后对 loss 的标量值调用.backward()来计算 Block 参数的 gradient。
理解 nn.Sequential 的内部机制
Section titled “理解 nn.Sequential 的内部机制”我们可以实现一个简化版的 Sequential 来理解其核心逻辑:
class MySequential(nn.Block): def __init__(self, **kwargs): super(MySequential, self).__init__(**kwargs) # self._children is an OrderedDict in the actual nn.Block to store sub-blocks
def add(self, block): # In a real Block, self._children stores blocks with unique names. # For simplicity, we use a list here. Gluon uses an OrderedDict. if not hasattr(self, '_added_blocks'): self._added_blocks = [] self._added_blocks.append(block) self.register_child(block) # Important for parameter collection etc.
def forward(self, x): for block in self._added_blocks: # Iterate through added blocks x = block(x) return x
# Using MySequentialx_custom = nd.random.uniform(shape=(2, 20), ctx=mx.cpu())N_net_custom = MySequential()with N_net_custom.name_scope(): N_net_custom.add(nn.Dense(256, activation='relu')) N_net_custom.add(nn.Dense(10))
N_net_custom.initialize(ctx=mx.cpu())output_custom = N_net_custom(x_custom)print(f"Output from MySequential (should be similar shape and type):\n{output_custom}")输出(实际值因单独初始化而不同):
<html> <body> <p>Output from MySequential (should be similar shape and type):[[ 0.01903114 0.02831168 0.01545286 -0.01109749 0.0350692 -0.02966933 -0.01100099 -0.02197874 -0.01193464 0.00189657] [ 0.0183696 0.02339937 0.01078339 -0.0076901 0.03513341 -0.02404617 -0.00755438 -0.01889096 -0.00740707 0.00199449]]<NDArray 2x10 @cpu(0)> </p> </body></html>自定义 Block
Section titled “自定义 Block”对于层不简单地按顺序排列的更复杂架构(例如,带有 skip connection 的网络如 ResNet,或多个输入/输出分支),您需要通过继承 gluon.Block(或 gluon.HybridBlock)来创建自定义 Block。
在自定义 Block 中,您通常需要:
- 在
__init__方法中将层定义为属性。 - 实现
forward(self, x, ...)方法来定义计算逻辑,使用在__init__中定义的层。
class MLP(nn.Block): def __init__(self, **kwargs): super(MLP, self).__init__(**kwargs) # Define layers within the name_scope to manage parameter names with self.name_scope(): self.hidden = nn.Dense(256, activation='relu') # Hidden layer self.output_layer = nn.Dense(10) # Output layer (renamed to avoid conflict with 'output' variable)
def forward(self, x): # Define the forward computation hidden_activation = self.hidden(x) final_output = self.output_layer(hidden_activation) return final_output
x_mlp = nd.random.uniform(shape=(2, 20), ctx=mx.cpu())N_net_mlp = MLP()N_net_mlp.initialize(ctx=mx.cpu())output_mlp = N_net_mlp(x_mlp)print(f"Output from custom MLP block:\n{output_mlp}")输出(值取决于初始化):
<html> <body> <p>Output from custom MLP block:[[-0.02450891 -0.02139075 -0.04023137 0.01311899 0.00839001 0.01907403 0.00343563 0.00571267 -0.02951591 -0.03287996] [-0.01569841 -0.01874819 -0.0320272 0.0182126 0.00212092 0.01575342 0.0099024 0.00390883 -0.02429737 -0.02714972]]<NDArray 2x10 @cpu(0)> </p> </body></html>虽然 Gluon 在 gluon.nn 中提供了许多常用层,但您可能需要从头开始实现一个自定义层。这可以通过继承 gluon.Block(用于纯命令式层)或 gluon.HybridBlock(用于可以被 hybridized 的层)来实现。
一个简单的自定义层(不含可学习参数)
Section titled “一个简单的自定义层(不含可学习参数)”自定义层需要实现 forward 方法。如果它没有可学习参数,__init__ 可能只需要调用父类构造函数。
# A layer that normalizes input to the range [0, 1]class NormalizationLayer(gluon.Block): def __init__(self, **kwargs): super(NormalizationLayer, self).__init__(**kwargs)
def forward(self, x): # Ensure min_val is not equal to max_val to avoid division by zero min_val = nd.min(x) max_val = nd.max(x) if min_val == max_val: return nd.zeros_like(x) return (x - min_val) / (max_val - min_val)
x_norm_layer_input = nd.array([[1, 2, 3], [4, 5, 6]], ctx=mx.cpu(), dtype='float32')norm_layer = NormalizationLayer()# No .initialize() needed if there are no parameters, but good practice if it might gain params later.# norm_layer.initialize(ctx=mx.cpu()) # Would be harmlessoutput_norm_layer = norm_layer(x_norm_layer_input)print(f"Output from NormalizationLayer:\n{output_norm_layer}")输出:
<html> <body> <p>Output from NormalizationLayer:[[0. 0.2 0.4] [0.6 0.8 1. ]]<NDArray 2x3 @cpu(0)> </p> </body></html>Hybridization:结合命令式和符号式执行
Section titled “Hybridization:结合命令式和符号式执行”gluon.HybridBlock 是一种特殊的 Block,它可以在两种模式下运行:命令式(像常规 Block 一样)和符号式。定义 HybridBlock 后,您可以调用其 .hybridize() 方法。这将尝试将 hybrid_forward 方法转换为符号图,然后 MXNet 可以对其进行优化以实现更快的执行。
hybridization 的好处:
- 灵活性:以命令式方式开发和调试。
- 性能:通过符号图优化进行部署,以提高速度并降低内存使用。
- 可移植性:导出符号图以便在其他语言绑定或设备上进行推理。
创建 HybridBlock 时,您需要实现 hybrid_forward(self, F, x, ...) 而不是 forward。F 参数至关重要:它引用 mxnet.ndarray(在命令式模式下运行)或 mxnet.symbol(在调用 .hybridize() 后构建符号图)。所有后端操作都必须使用 F(例如,F.relu(x) 而不是 nd.relu(x))。
示例:一个 Hybrid 归一化层
Section titled “示例:一个 Hybrid 归一化层”class NormalizationHybridLayer(gluon.HybridBlock): def __init__(self, **kwargs): super(NormalizationHybridLayer, self).__init__(**kwargs)
def hybrid_forward(self, F, x): min_val = F.min(x) max_val = F.max(x) # Add a small epsilon to prevent division by zero if min_val == max_val # F.broadcast_xxx operations are common for N-D array compatibility denominator = F.broadcast_sub(max_val, min_val) + 1e-8 return F.broadcast_div(F.broadcast_sub(x, min_val), denominator)
input_hybrid_norm = nd.array([1, 2, 3, 4, 5, 6], ctx=mx.cpu(), dtype='float32')layer_hybd = NormalizationHybridLayer()# layer_hybd.initialize() # Not strictly needed without params, but good practice
# Run imperatively firstoutput_imperative = layer_hybd(input_hybrid_norm)print(f"Imperative output:\n{output_imperative}")
# Hybridize the layerlayer_hybd.hybridize()
# Run again (now using symbolic graph internally)output_hybridized = layer_hybd(input_hybrid_norm)print(f"Hybridized output:\n{output_hybridized}")输出:
<html> <body> <p>Imperative output:[0. 0.2 0.4 0.6 0.8 1. ]<NDArray 6 @cpu(0)>Hybridized output:[0. 0.2 0.4 0.6 0.8 1. ]<NDArray 6 @cpu(0)> </p> </body></html>注意:Hybridization 是将模型编译成符号图的一个步骤。执行时使用 CPU 还是 GPU 是独立于此的,取决于输入数据和层参数的 context。
Block 和 HybridBlock 的区别
Section titled “Block 和 HybridBlock 的区别”主要区别在于是否可以 hybridization。Block 始终以命令式方式运行。HybridBlock 在调用 .hybridize() 之前可以以命令式方式运行,之后则以符号式(编译后)方式运行。这要求在 hybrid_forward 中使用 F 参数进行所有 MXNet 操作,以确保与 mx.nd 和 mx.sym 后端兼容。
在网络中使用自定义层
Section titled “在网络中使用自定义层”自定义层,特别是 HybridBlock,可以使用像 gluon.nn.Sequential 或 gluon.nn.HybridSequential 这样的容器与标准层结合使用。HybridSequential 本身就是一个 HybridBlock,用于堆叠其他 HybridBlock。
net = gluon.nn.HybridSequential()with net.name_scope(): net.add(nn.Dense(5, activation='relu')) # Standard dense layer net.add(NormalizationHybridLayer()) # Our custom hybrid layer net.add(nn.Dense(1)) # Another standard dense layer
net.initialize(mx.init.Xavier(magnitude=2.24), ctx=mx.cpu())
# Hybridize the entire networknet.hybridize()
# Dummy inputinput_data_for_net = nd.random.uniform(low=-10, high=10, shape=(10, 2), ctx=mx.cpu())output_from_net = net(input_data_for_net)print(f"Output from hybrid sequential net (shape {output_from_net.shape}):\n{output_from_net[0:3]}") # Print first 3 rows输出(前 3 行,实际值会有所不同):
<html> <body> <p>Output from hybrid sequential net (shape (10, 1)):[[-0.40172195] [-0.40172195] [-0.40172195]]<NDArray 3x1 @cpu(0)> </p> </body></html>自定义层参数
Section titled “自定义层参数”层通常具有可学习参数(weights, biases)或固定的配置参数。在 Gluon 中,参数通过 self.params 管理,它是每个 Block 中可用的 ParameterDict 对象。
- 可学习参数:使用
self.params.get(...)定义。它们的 shape 通常可以被 inferred (allow_deferred_init=True)。它们在训练期间会被更新。 - 固定参数/Buffer:如果需要,也可以存储,有时作为常量或不可微分参数。
示例:带有可学习 Weights 和固定 Scales 的自定义层
Section titled “示例:带有可学习 Weights 和固定 Scales 的自定义层”class ScaledFullyConnectedLayer(gluon.HybridBlock): def __init__(self, hidden_units, fixed_scales_data, **kwargs): super(ScaledFullyConnectedLayer, self).__init__(**kwargs) self.hidden_units = hidden_units with self.name_scope(): # Learnable weights for a fully connected operation # Shape is (output_units, input_units_or_0_for_inference) self.weights = self.params.get('weights', shape=(hidden_units, 0), # Input dim inferred later allow_deferred_init=True, init=mx.init.Normal(0.01)) # Fixed scales (not trainable) self.scales = self.params.get('scales', shape=fixed_scales_data.shape, init=mx.init.Constant(fixed_scales_data.asnumpy()), differentiable=False) # Not updated by optimizer # Bias (learnable) self.bias = self.params.get('bias', shape=(hidden_units,), init=mx.init.Constant(0), allow_deferred_init=True)
def hybrid_forward(self, F, x, weights, scales, bias): # Fully connected operation using F for backend compatibility fc_out = F.FullyConnected(data=x, weight=weights, bias=bias, num_hidden=self.hidden_units, no_bias=False) # Element-wise multiplication with scales scaled_out = F.broadcast_mul(fc_out, scales) return scaled_out
# Example usagenum_features_in = 10num_hidden_units = 5fixed_scales = nd.array([0.1, 0.5, 1.0, 0.5, 0.1], ctx=mx.cpu()) # Scales for each hidden unit
scaled_fc_layer = ScaledFullyConnectedLayer(hidden_units=num_hidden_units, fixed_scales_data=fixed_scales)scaled_fc_layer.initialize(ctx=mx.cpu())
dummy_input_fc = nd.random.normal(shape=(3, num_features_in), ctx=mx.cpu())output_scaled_fc = scaled_fc_layer(dummy_input_fc)
print(f"Input shape: {dummy_input_fc.shape}")print(f"Output shape from ScaledFullyConnectedLayer: {output_scaled_fc.shape}")print(f"Parameters in layer: {scaled_fc_layer.collect_params()}")