Skip to content

Apache MXNet - NDArray

在本章中,我们将探讨 NDArray,它是 Apache MXNet 处理多维数组格式的数值数据的主要工具,类似于 NumPy 数组,但增加了深度学习和 GPU 计算的功能。

让我们首先了解如何使用 NDArray 创建和操作数据。您需要满足以下先决条件:

为了有效地使用 NDArray 并跟随本章的示例,请确保您已具备以下条件:

  • 在您的 Python 环境中安装了 Apache MXNet(推荐版本 1.6 或更高)。通常可以使用 pip 安装:pip install mxnet(或根据您的 CUDA 版本选择带有 GPU 支持的变体,如 pip install mxnet-cu118)。
  • Python 3.7 或更高版本。

让我们深入一些示例,了解基本的 NDArray 功能。

首先,导入 MXNet 库及其 NDArray 模块。一种常见的约定是将 ndarray 导入为 nd:

import mxnet as mx
from mxnet import nd

现在,让我们探索各种创建 NDArray 的方法:

示例

x = nd.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10])
print(x)

输出

输出显示一个 1-D array:

[ 1. 2. 3. 4. 5. 6. 7. 8. 9. 10.]
<ndarray shape=(10,) cpu(0)>

示例

y = nd.array([
[1, 2, 3, 4, 5],
[6, 7, 8, 9, 10],
[11, 12, 13, 14, 15]
])
print(y)

输出

输出显示一个 2-D array (matrix):

[[ 1. 2. 3. 4. 5.]
[ 6. 7. 8. 9. 10.]
[11. 12. 13. 14. 15.]]
<ndarray shape=(3, 5) cpu(0)>

使用特定的 Initializer 创建 NDArray

Section titled “使用特定的 Initializer 创建 NDArray”

MXNet 提供了函数来创建具有指定 shape 和初始内容的 NDArray:

示例:使用 nd.empty, nd.full

nd.empty(shape, ctx) 在给定的 context 上创建一个指定 shape 的 array,但其条目未初始化(包含内存中的任意数据)。如果您计划立即填充它,它会稍微快一些,但请谨慎使用。

nd.full(shape, value, ctx) 在给定的 shape 和 context 上创建一个 array,并用指定的 scalar value 完全填充。

a = nd.empty((2, 3), ctx=mx.cpu()) # Uninitialized 2x3 array on CPU
print("Uninitialized array (a):")
print(a)
b = nd.full((2, 3), 8.0, ctx=mx.cpu()) # 2x3 array filled with 8.0 on CPU
print("\nArray filled with 8.0 (b):")
print(b)

输出

输出将类似于此(nd.empty 的确切值每次可能不同):

Uninitialized array (a):
[[0.000e+00 0.000e+00 0.000e+00]
[0.000e+00 2.887e-42 0.000e+00]]
<ndarray shape=(2, 3) cpu(0)>
Array filled with 8.0 (b):
[[8. 8. 8.]
[8. 8. 8.]]
<ndarray shape=(2, 3) cpu(0)>

使用 nd.zeros(shape, ctx) 在指定的 context 上创建一个全零 array。

示例

z = nd.zeros((3, 4), ctx=mx.cpu())
print(z)

输出

输出是一个 3x4 的全零 matrix:

[[0. 0. 0. 0.]
[0. 0. 0. 0.]
[0. 0. 0. 0.]]
<ndarray shape=(3, 4) cpu(0)>

使用 nd.ones(shape, ctx) 在指定的 context 上创建一个全一 array。

示例

o = nd.ones((2, 5), ctx=mx.cpu())
print(o)

输出

输出是一个 2x5 的全一 matrix:

[[1. 1. 1. 1. 1.]
[1. 1. 1. 1. 1.]]
<ndarray shape=(2, 5) cpu(0)>

MXNet 的 nd.random 子模块提供了各种 distributions。例如,nd.random.normal(loc, scale, shape, ctx) 在指定的 context 上从 normal (Gaussian) distribution 中采样。

示例

# Sample from a normal distribution with mean 0 and standard deviation 1
r = nd.random.normal(loc=0, scale=1, shape=(3, 4), ctx=mx.cpu())
print(r)

输出

输出将是一个 3x4 的随机数 array(每次运行值都会变化):

[[ 1.2673576 -2.0345826 -0.32537818 -1.4583491 ]
[-0.11176403 1.3606371 -0.7889914 -0.17639421]
[-0.2532185 -0.42614475 -0.12548696 1.4022992 ]]
<ndarray shape=(3, 4) cpu(0)>

您可以检查 NDArray 的关键属性:

示例:获取上面创建的随机 array r 的 shape、size、data type 和 context。

print("Shape of r:", r.shape)
print("Size of r (total number of elements):", r.size)
print("Data type of r:", r.dtype)
print("Context of r:", r.context)

输出

输出将是:

Shape of r: (3, 4)
Size of r (total number of elements): 12
Data type of r: <class 'numpy.float32'>
Context of r: cpu(0)

注意:MXNet 默认使用 32 位 floating-point numbers (float32)。由于 MXNet 的 interoperability,dtype 通常报告为 NumPy dtype。.context 属性显示 array 的数据驻留在何处(例如,cpu(0) 表示第一个 CPU,gpu(0) 表示第一个 GPU)。

NDArray 支持一套全面的数学操作,其中许多操作类似于 NumPy。除非另有指定,这些操作通常是 element-wise 的,并且是 context-aware 的(即,操作通常发生在输入 array 的 context 上)。

示例

import mxnet as mx
from mxnet import nd
# Ensure arrays are on the same context for operations
ctx = mx.cpu()
x = nd.ones((2, 3), ctx=ctx)
y = nd.arange(6, ctx=ctx).reshape((2, 3)) # Creates [[0,1,2],[3,4,5]]
print('x =\n', x)
print('y =\n', y)
result = x + y
print('result (x + y) =\n', result)

输出

输出展示了 element-wise addition:

x =
[[1. 1. 1.]
[1. 1. 1.]]
<ndarray shape=(2, 3) cpu(0)>
y =
[[0. 1. 2.]
[3. 4. 5.]]
<ndarray shape=(2, 3) cpu(0)>
result (x + y) =
[[1. 2. 3.]
[4. 5. 6.]]
<ndarray shape=(2, 3) cpu(0)>

示例

ctx = mx.cpu()
a = nd.array([1, 2, 3, 4], ctx=ctx)
b = nd.array([2, 2, 2, 1], ctx=ctx)
product = a * b
print('product (a * b) =', product)

输出

输出是:

product (a * b) = [2. 4. 6. 4.]
<ndarray shape=(4,) cpu(0)>

nd.exp(array) 计算 element-wise exponential (e^x)。

示例

c = nd.array([1, 2, 3], ctx=mx.cpu())
exp_c = nd.exp(c)
print('exp_c =', exp_c)

输出

输出显示 e 的每个元素的幂:

exp_c = [ 2.7182817 7.389056 20.085537 ]
<ndarray shape=(3,) cpu(0)>

使用 nd.dot(A, B) 进行 matrix multiplication。对于 2-D array,这将计算标准的 matrix product。Python 3.5+ 也支持 @ operator 进行 matrix multiplication,因其简洁性而常被优先使用。

示例

ctx = mx.cpu()
mat1 = nd.arange(6, ctx=ctx).reshape((2, 3)) # 2x3 matrix: [[0,1,2],[3,4,5]]
mat2 = nd.arange(3, ctx=ctx).reshape((3, 1)) # 3x1 matrix: [[0],[1],[2]]
# Using nd.dot
dot_product = nd.dot(mat1, mat2)
print('Using nd.dot:\n', dot_product)
# Using @ operator (requires Python 3.5+)
at_product = mat1 @ mat2
print('\nUsing @ operator:\n', at_product)

输出

两种方法都产生相同的 2x1 matrix:

Using nd.dot:
[[ 5.]
[14.]]
<ndarray shape=(2, 1) cpu(0)>
Using @ operator:
[[ 5.]
[14.]]
<ndarray shape=(2, 1) cpu(0)>

标准操作如 y = x + y 为结果 y 分配新的内存。在深度学习中,这对于大型 array 或频繁更新(例如,训练期间的模型 parameter)来说可能效率低下。MXNet 支持 in-place operations 来修改现有的 array 内存,从而减少 overhead。

让我们说明内存分配。Python 的 id() 函数返回一个对象的唯一标识符(在 CPython 中通常是其内存地址)。

ctx = mx.cpu()
x = nd.arange(4, ctx=ctx)
y = nd.ones(4, ctx=ctx)
print('Original y =', y)
print('id(y) before addition:', id(y))
y_new = y + x # Standard addition allocates new memory for y_new
print('\ny_new after y_new = y + x:', y_new)
print('id(y_new) after addition (new memory):', id(y_new))
print('id(y) (original y is unchanged):', id(y))

输出

注意 id(y_new) 与 id(y) 的不同,表明 y_new 指向新的内存位置:

Original y = [1. 1. 1. 1.]
<ndarray shape=(4,) cpu(0)>
id(y) before addition: 140123456789012
y_new after y_new = y + x: [1. 2. 3. 4.]
<ndarray shape=(4,) cpu(0)>
id(y_new) after addition (new memory): 140123456789134
id(y) (original y is unchanged): 140123456789012

(注意:确切的 id 值会有所不同。)

要执行 in-place 操作,可以使用 slice assignment 或带有 out 参数的函数:

  1. 使用 slice assignment [:]:
ctx = mx.cpu()
x = nd.arange(4, ctx=ctx)
y_orig = nd.ones(4, ctx=ctx)
z = nd.zeros_like(x) # Pre-allocate z on the same context as x
print('id(z) before:', id(z))
# Result of y_orig + x is first computed (possibly in new memory),
# then copied into z's existing memory via slice assignment.
z[:] = y_orig + x
print('\nz after z[:] = y_orig + x:', z)
print('id(z) after (should be same):', id(z))

输出

id(z) 保持不变,但在复制之前可能使用了临时 buffer 来存储 y_orig + x 的结果。

id(z) before: 140123456789256
z after z[:] = y_orig + x: [1. 2. 3. 4.]
<ndarray shape=(4,) cpu(0)>
id(z) after (should be same): 140123456789256
  1. 使用 out 参数(最节省内存):许多 MXNet operator(如 nd.add, nd.multiply)接受一个 out 参数来指定目标 array。这避免了操作本身的临时分配。
ctx = mx.cpu()
x = nd.arange(4, ctx=ctx)
y = nd.ones(4, ctx=ctx)
z_out = nd.zeros_like(x) # Pre-allocate z_out
print('id(z_out) before:', id(z_out))
nd.add(x, y, out=z_out) # Efficient in-place capable addition
print('\nz_out after nd.add(x, y, out=z_out):', z_out)
print('id(z_out) after (should be same):', id(z_out))

输出

id(z_out) 保持不变,理想情况下,加法操作本身不需要中间 buffer。

id(z_out) before: 140123456789378
z_out after nd.add(x, y, out=z_out): [1. 2. 3. 4.]
<ndarray shape=(4,) cpu(0)>
id(z_out) after (should be same): 140123456789378

诸如 x += y 或 x *= y 之类的 in-place 缩写 operator 在支持的情况下也会直接修改 x。例如,z_out += x(在 z_out 包含 x+y 之后)会在 z_out 的内存可以重用时将 x in-place 添加到 z_out。

MXNet 允许您为 NDArray 指定计算 context (CPU 或 GPU)。这对于性能至关重要,因为深度学习任务在 GPU 上的计算通常快得多。默认情况下,除非设置了不同的默认 context (mx.default_context()),否则 NDArray 会在 CPU 上创建。

您可以检查 GPU 是否可用:num_gpus = mx.context.num_gpus()。如果 num_gpus 为 0,则 GPU 示例将引发错误或需要修改。

示例:在特定的 CPU 上创建一个 array,然后复制到 GPU(如果可用)。

import mxnet as mx
from mxnet import nd
# Create an array on the default CPU (often cpu(0))
default_cpu_ctx = mx.cpu() # Or mx.cpu(0)
a_cpu = nd.ones(shape=(3, 3), ctx=default_cpu_ctx)
print("Array on CPU:")
print(a_cpu)
# Check if GPUs are available
num_gpus = mx.context.num_gpus()
print(f"\nNumber of GPUs available: {num_gpus}")
if num_gpus > 0:
gpu_ctx = mx.gpu(0) # Use the first GPU
# Copy array to the GPU
a_gpu = a_cpu.copyto(gpu_ctx)
print("\nArray copied to GPU:")
print(a_gpu)
# Alternative way to get an array on a context:
# as_in_context returns the array itself if already on target context,
# otherwise, it copies and returns a new array on the target context.
b_gpu = nd.ones(shape=(2,2), ctx=gpu_ctx)
c_gpu = a_cpu.as_in_context(gpu_ctx) # c_gpu is a new array on gpu(0)
d_gpu = b_gpu.as_in_context(gpu_ctx) # d_gpu is b_gpu itself (no copy)
print("\nArray c_gpu (from as_in_context on a_cpu):")
print(c_gpu)
print(f"id(b_gpu) == id(d_gpu): {id(b_gpu) == id(d_gpu)}")
else:
print("\nNo GPUs available. Skipping GPU-specific copy examples.")

输出(如果存在 GPU):

Array on CPU:
[[1. 1. 1.]
[1. 1. 1.]
[1. 1. 1.]]
<ndarray shape=(3, 3) cpu(0)>
Number of GPUs available: 1 (output may vary)
Array copied to GPU:
[[1. 1. 1.]
[1. 1. 1.]
[1. 1. 1.]]
<ndarray shape=(3, 3) gpu(0)>
Array c_gpu (from as_in_context on a_cpu):
[[1. 1. 1.]
[1. 1. 1.]
[1. 1. 1.]]
<ndarray shape=(3, 3) gpu(0)>
id(b_gpu) == id(d_gpu): True

涉及不同 contexts 上的 NDArray 的操作(例如,将 CPU array 添加到 GPU array)通常不允许直接进行。数据必须显式移动到同一 context。在 CPU 和 GPU 之间传输数据会产生 overhead,因此在性能关键的代码中最好尽量减少这些传输。

虽然 NDArray 设计类似于 NumPy 的 array,但两者之间存在一些关键区别,主要与深度学习优化有关:

  • GPU Acceleration: NDArray 原生支持 GPU 计算,这对于高效训练深度神经网络至关重要。
  • Asynchronous Execution: MXNet 在 NDArray 上的操作通常是 asynchronously 执行的。当您发出 c = a + b 这样的命令时,操作被推送到 MXNet 的 execution engine。Python interpreter 可能会立即重新获得控制权,即使计算尚未完成。Engine 管理 dependencies 并优化执行顺序。这允许 parallelism 和 efficiency,尤其是在 GPU 上。实际结果通常只有在明确请求时(例如,打印、通过 .asnumpy() 转换为 NumPy、或保存到磁盘)才会计算并可用。
  • Automatic Differentiation: NDArray 与 MXNet 的 autograd 模块无缝集成,用于 automatic calculation of gradients (derivatives),这对于通过 backpropagation 训练神经网络是基础。

要确保计算完成并获取其实际结果,您可以在 NDArray 上调用 .wait_to_read()。使用 .asnumpy() 将其转换为 NumPy array 也是一个 blocking call,会等待完成。

MXNet 提供了 NDArray 和 NumPy array 之间的无缝转换,方便与更广泛的 Python 科学计算生态系统集成。

NDArray 到 NumPy: 使用 .asnumpy() 方法。

import numpy as np
from mxnet import nd
nd_arr = nd.array([1, 2, 3], ctx=mx.cpu())
np_arr = nd_arr.asnumpy() # Convert NDArray to NumPy array
print("NDArray:", nd_arr)
print("NumPy array:", np_arr)
print("Type of np_arr:", type(np_arr))

输出:

NDArray: [1. 2. 3.]
<ndarray shape=(3,) cpu(0)>
NumPy array: [1. 2. 3.]
Type of np_arr: <class 'numpy.ndarray'>

NumPy 到 NDArray: 使用 nd.array(numpy_array, ctx)。

np_data = np.array([[4, 5], [6, 7]])
nd_data = nd.array(np_data, ctx=mx.cpu()) # Convert NumPy array to NDArray on CPU
print("Original NumPy array:\n", np_data)
print("Converted NDArray:\n", nd_data)

输出:

Original NumPy array:
[[4 5]
[6 7]]
Converted NDArray:
[[4. 5.]
[6. 7.]]
<ndarray shape=(2, 2) cpu(0)>

重要:将位于 GPU 上的 NDArray 使用 .asnumpy() 转换为 NumPy array 会隐式地将数据从 GPU 复制到 CPU 内存。此操作可能很慢,在性能关键的代码段中应谨慎使用。

您通常可以通过组合现有的 NDArray operator 来实现与某些 NumPy 函数类似的功能。例如,np.full_like(a, val)(创建一个与 a 具有相同 shape 并填充 val 的 array)可以在 MXNet 中复制:

a_np = np.arange(5, dtype=int)
np_x = np.full_like(a=a_np, fill_value=10)
# MXNet equivalent
ctx = mx.cpu()
a_nd = nd.arange(5, ctx=ctx)
# Create an array of ones with the same shape and context as a_nd, then multiply by 10
nd_x = nd.ones_like(a_nd) * 10
# Alternatively, more directly:
# nd_x_alt = nd.full(a_nd.shape, 10, ctx=a_nd.context, dtype=a_nd.dtype)
print("NumPy full_like result:", np_x)
print("MXNet equivalent result:", nd_x)
print("Are they equal (after converting nd_x to numpy)?", np.array_equal(np_x, nd_x.asnumpy()))

输出:

NumPy full_like result: [10 10 10 10 10]
MXNet equivalent result: [10. 10. 10. 10. 10.]
<ndarray shape=(5,) cpu(0)>
Are they equal (after converting nd_x to numpy)? True

虽然许多 MXNet NDArray 函数模仿 NumPy,但名称或签名可能略有不同(例如,nd.split() 对比 np.split())。请始终查阅官方 Apache MXNet API 文档以获取精确用法。为了在 MXNet 中获得更直接的 NumPy-like 体验,可以考虑使用 mxnet.numpy 模块(通常导入为 mx.np),它提供的 API 紧密遵循 NumPy 的 conventions,并通常提供更好的 NumPy 兼容性。

在 Training Loops 中最小化 Blocking Calls 的影响

Section titled “在 Training Loops 中最小化 Blocking Calls 的影响”

像 .asnumpy() 或 .asscalar()(用于从单元素 NDArray 获取 Python scalar)这样的 calls 是 blocking 的:它们强制 MXNet 完成该 array 的所有挂起计算,并在必要时将数据从其当前 context(例如 GPU)复制到 CPU 内存。在 training loops 中,频繁的 blocking calls 会通过 serializing execution 产生性能瓶颈。

缓解这种情况的一种策略是延迟这些 blocking calls,直到真正需要结果时,或者理想情况下,将它们与其他独立计算重叠。例如,在 logging 训练 loss 时,不是在每次 iteration 中立即将 loss 转换为 NumPy array,而是可以缓冲 NDArray loss,并转换/打印 上一个 iteration 的 loss。这使得当前 iteration 的计算(包括 GPU 上的 backpropagation)有可能与上一步的(blocking、CPU-bound).asnumpy() call 重叠。

以下示例使用简单的 MXNet Gluon API training loop 演示了这个概念。LossBuffer 类有助于延迟对 loss 值进行 .asnumpy() call。

import mxnet as mx
from mxnet import gluon, nd, autograd
import numpy as np
import time # For demonstration
class LossBuffer(object):
"""Simple buffer for storing and retrieving loss, delaying asnumpy()."""
def __init__(self):
self._loss_nd_current_iter = None # Stores NDArray loss of current iteration
self._loss_np_previous_iter = None # Stores NumPy loss from previous iteration
def log_and_buffer(self, new_loss_nd):
"""Converts previous loss to NumPy, buffers new loss as NDArray, returns previous NumPy loss."""
# Convert previous iteration's loss (if it exists) to NumPy
if self._loss_nd_current_iter is not None:
# This asnumpy() call is on the *previous* iteration's loss.
# It might overlap with the current iteration's GPU work.
self._loss_np_previous_iter = np.mean(self._loss_nd_current_iter.asnumpy())
else:
self._loss_np_previous_iter = None # No previous loss on first call
self._loss_nd_current_iter = new_loss_nd # Store current loss (still an NDArray)
return self._loss_np_previous_iter
# Setup: A simple model, loss, and data
ctx = mx.gpu() if mx.context.num_gpus() > 0 else mx.cpu() # Use GPU if available, else CPU
print(f"Training on: {ctx}")
net = gluon.nn.Dense(10, prefix='dense_') # Output layer for 10 classes
net.initialize(ctx=ctx) # Initialize parameters on the chosen context
loss_fn = gluon.loss.SoftmaxCrossEntropyLoss() # Standard loss for classification
# Dummy data parameters
num_samples = 1024
num_features = 100
num_batches = 8
batch_size = num_samples // num_batches
num_epochs = 1 # For a quick demo
# Create dummy data directly on the target context
data = nd.random.uniform(shape=(num_samples, num_features), ctx=ctx)
label = nd.array(np.random.randint(0, 10, (num_samples,)), ctx=ctx, dtype='int32')
train_dataset = gluon.data.ArrayDataset(data, label)
# Using num_workers > 0 can speed up data loading by using separate processes.
# However, for simplicity or if issues arise, num_workers=0 (loads in main process) is safer.
train_data_loader = gluon.data.DataLoader(train_dataset, batch_size=batch_size, shuffle=True, num_workers=0)
trainer = gluon.Trainer(net.collect_params(), optimizer='sgd', optimizer_params={'learning_rate': 0.01})
loss_buffer = LossBuffer()
for epoch in range(num_epochs):
epoch_start_time = time.time()
cumulative_loss_sum_for_epoch = 0
num_samples_in_epoch = 0
for i, (batch_data, batch_label) in enumerate(train_data_loader):
# Data might be on CPU from DataLoader if num_workers > 0; ensure it's on target context
batch_data = batch_data.as_in_context(ctx)
batch_label = batch_label.as_in_context(ctx)
with autograd.record(): # Start recording computations for gradient calculation
output = net(batch_data)
current_batch_loss_nd = loss_fn(output, batch_label) # Loss is an NDArray on ctx
# This call saves the new NDArray loss and returns the *previous* iteration's loss as NumPy
previous_avg_loss_np = loss_buffer.log_and_buffer(current_batch_loss_nd)
current_batch_loss_nd.backward() # Backpropagate based on current_batch_loss_nd
trainer.step(batch_data.shape[0]) # Update model parameters
if previous_avg_loss_np is not None:
# Print the average loss from the *previous* iteration (already converted)
print(f"Epoch {epoch}, Batch {i-1}: Prev. Iter. Avg Batch Loss: {previous_avg_loss_np:.4f}")
# Accumulate current batch's loss for epoch average (on NDArray to avoid blocking)
# Note: current_batch_loss_nd is typically a vector of losses, one per sample.
# We sum it here; it will be averaged at the end of epoch.
cumulative_loss_sum_for_epoch += current_batch_loss_nd.sum()
num_samples_in_epoch += batch_data.shape[0]
# After the loop, process the final buffered loss (from the last batch of the epoch)
final_buffered_avg_loss_np = loss_buffer.log_and_buffer(None) # Clear buffer, get last batch's loss
if final_buffered_avg_loss_np is not None:
print(f"Epoch {epoch}, Batch {i}: Prev. Iter. Avg Batch Loss: {final_buffered_avg_loss_np:.4f}")
# Calculate and print average epoch loss (now we need to block for sum)
avg_epoch_loss = cumulative_loss_sum_for_epoch.asscalar() / num_samples_in_epoch
print(f"Epoch {epoch} completed. Avg Epoch Loss: {avg_epoch_loss:.4f}. Time: {time.time() - epoch_start_time:.2f}s.")

输出(因随机数据/初始化和 context 而异;loss 值仅为示例):

Training on: cpu(0) (or gpu(0) if available)
Epoch 0, Batch 0: Prev. Iter. Avg Batch Loss: 2.3512
Epoch 0, Batch 1: Prev. Iter. Avg Batch Loss: 2.3488
...
Epoch 0, Batch 6: Prev. Iter. Avg Batch Loss: 2.3078
Epoch 0, Batch 7: Prev. Iter. Avg Batch Loss: 2.3002
Epoch 0 completed. Avg Epoch Loss: 2.3250. Time: 0.15s.

这种交错进行 .asnumpy() call 的技术有助于将这种潜在的 blocking CPU-bound 任务与后续 iteration 的独立 GPU 计算重叠,从而提高整体 training throughput。有关 MXNet 中更高级的 profiling 和性能优化,请参阅官方 Apache MXNet 文档和 MXNet profiler 等工具。有关 NDArray 操作和 best practices 的更多资源也可在其中找到。