位置编码
理解 Transformer 模型中的位置编码(Positional Encoding)
Section titled “理解 Transformer 模型中的位置编码(Positional Encoding)”Transformer 模型首先将输入标记(token,可以是单词、子词或字符)转换为向量表示,称为输入嵌入(input embeddings)。然而,这些嵌入本身并不固有地包含序列中标记的顺序或位置信息。由于词序对于语言的意义至关重要(例如,“狗咬人”与“人咬狗”),Transformer 需要一种机制来整合这种位置信息。这就是通过位置编码(Positional Encoding)实现的。
位置编码被添加到编码器(encoder)和解码器(decoder)堆栈底部的输入嵌入中。本章解释了位置编码是什么、为什么它是必需的、一种常见方法如何工作,并提供了一个概念性的 Python 实现。
什么是位置编码?
Section titled “什么是位置编码?”位置编码是一种用于 Transformer 模型的技术,旨在注入序列中标记的相对或绝对位置信息。由于 Transformer 中的自注意力机制(self-attention mechanism)并行处理所有标记,它本身并不考虑它们的顺序。位置编码提供了这种缺失的序列上下文。
在原始的 Transformer 架构中,位置编码向量与输入嵌入的维度相同(d_model),这使得它们可以直接相加。这种组合表示(嵌入 + 位置编码)随后作为输入传递给 Transformer 的后续层。
位置编码组件通常位于 Transformer 的编码器和解码器部分的输入嵌入子层之后。输入嵌入提供标记的语义含义,而位置编码则提供它们在序列中的上下文。
为什么 Transformer 中需要位置编码?
Section titled “为什么 Transformer 中需要位置编码?”与循环神经网络(RNNs)或长短期记忆网络(LSTMs)不同,后者逐个处理序列中的标记,因此固有地包含顺序信息,Transformer 同时处理序列中的所有标记。自注意力机制计算所有标记对之间的注意力分数,无论它们之间的距离如何。如果没有位置信息,Transformer 将把句子视为一个“词袋”,从而失去关键的序列意义。
例如,句子“猫追狗”和“狗追猫”使用相同的词语,但由于词序不同而含义不同。位置编码通过使每个标记的表示基于其在序列中的位置而独一无二,从而确保模型能够区分这些情况。
正弦位置编码(Sinusoidal Positional Encoding)如何工作?
Section titled “正弦位置编码(Sinusoidal Positional Encoding)如何工作?”虽然位置信息可以通过学习获得(例如,为每个位置设置可学习的嵌入向量),但 Vaswani 等人(2017)在原始 Transformer 论文中提出使用固定的正弦函数。这种方法有几个优点:它可以泛化到比训练期间见过的更长的序列,并且它允许模型轻松学习关注相对位置,因为 PE(pos+k) 可以表示为 PE(pos) 的线性函数。
对于序列中位置为 ‘pos’ 且编码向量维度为 ‘i’ 的标记,位置编码 (PE) 定义如下,使用不同频率的正弦和余弦函数:
$$\mathrm{PE}(\text{pos}, 2i) = \sin\left(\frac{\text{pos}}{10000^{\frac{2i}{d_{\text{model}}}}}\right)$$
$$\mathrm{PE}(\text{pos}, 2i+1) = \cos\left(\frac{\text{pos}}{10000^{\frac{2i}{d_{\text{model}}}}}\right)$$
其中:
- pos 是标记在序列中的位置 (0, 1, 2, …)。
- i 是位置编码向量中的维度索引 (0, 1, …, d_model/2 - 1)。
- d_model 是嵌入的维度(因此也是位置编码向量的维度,例如 512)。
- 项 $10000^{\frac{2i}{d_{\text{model}}}}$ 创建了波长范围从 $2\pi$ 到 $10000 \cdot 2\pi$ 的函数。位置编码的每个维度对应一个正弦函数。对于偶数维度 (2i),使用正弦函数;对于奇数维度 (2i+1),使用余弦函数。
这种公式确保了每个位置都有唯一的编码。此外,选择正弦函数使得模型有可能外推到比训练时遇到的更长的序列。
正弦位置编码的概念性 Python 实现
Section titled “正弦位置编码的概念性 Python 实现”以下是一个使用 NumPy 生成正弦位置编码并将其添加到示例输入嵌入的 Python 脚本。这演示了核心概念。
import numpy as np
# --- Step 1: Example Text, Tokenization, and Input Embeddings ---# Example sentence and basic tokenizationtext = "Transformers use positional encoding"tokens = text.lower().split()
# Create a simple vocabularyvocab = {word: idx for idx, word in enumerate(tokens)}
# Example input: sequence of token indicesinput_indices = np.array([vocab[word] for word in tokens])
print(f"Vocabulary: {vocab}")print(f"Input Indices: {input_indices}")
# Parameters for embeddingsvocab_size = len(vocab)embed_dim = 6 # Using a small dimension for easier visualization (d_model)
# Initialize a dummy embedding matrix (e.g., learned during training)# In a real scenario, these would be meaningful learned vectors.np.random.seed(42) # for reproducibilityembedding_matrix = np.random.rand(vocab_size, embed_dim)
# Get the embeddings for the input indicesinput_embeddings = embedding_matrix[input_indices]
print(f"\nEmbedding Matrix (shape: {embedding_matrix.shape}):\n{embedding_matrix}")print(f"Input Embeddings (shape: {input_embeddings.shape}):\n{input_embeddings}")
# --- Step 2: Positional Encoding Function ---def get_positional_encoding(max_seq_len, d_model): """Generates sinusoidal positional encodings.
Args: max_seq_len: Maximum sequence length. d_model: Dimensionality of the model (embedding dimension).
Returns: A numpy array of shape (max_seq_len, d_model) with positional encodings. """ pe = np.zeros((max_seq_len, d_model)) position = np.arange(0, max_seq_len).reshape(-1, 1) # Column vector for positions
# Term for calculating frequencies/wavelengths # div_term for even dimensions (0, 2, 4, ...) div_term = np.exp(np.arange(0, d_model, 2) * -(np.log(10000.0) / d_model))
pe[:, 0::2] = np.sin(position * div_term) # Apply sin to even indices in pe pe[:, 1::2] = np.cos(position * div_term) # Apply cos to odd indices in pe
return pe
# --- Step 3: Generate and Apply Positional Encodings ---max_len_sequence = len(tokens) # Max length for this specific sequence
# Generate positional encodingspos_encodings = get_positional_encoding(max_len_sequence, embed_dim)
print(f"\nPositional Encodings (shape: {pos_encodings.shape}):\n{pos_encodings}")
# Add positional encodings to input embeddings# The `pos_encodings` matrix might be pre-calculated for a max possible length.# Here, we take only up to the current sequence length.# (In this example, max_len_sequence is already len(tokens), so pos_encodings[:len(tokens)] is just pos_encodings)final_embeddings = input_embeddings + pos_encodings
print(f"\nInput Embeddings with Positional Encoding (shape: {final_embeddings.shape}):\n{final_embeddings}")预期输出(示意)
Section titled “预期输出(示意)”运行脚本将产生输出,显示词汇表、输入索引、随机初始化的嵌入矩阵、从中派生的输入嵌入、生成的职位编码,最后是输入嵌入与位置编码的组合。还将打印矩阵的形状。(注意:除非固定了示例中的种子,否则随机嵌入的精确数值会因运行而异)。
Vocabulary: {'transformers': 0, 'use': 1, 'positional': 2, 'encoding': 3}Input Indices: [0 1 2 3]
Embedding Matrix (shape: (4, 6)):[[0.37454012 0.95071431 0.73199394 0.59865848 0.15601864 0.15599452] [0.05808361 0.86617615 0.60111501 0.70807258 0.02058449 0.96990985] [0.83244264 0.21233911 0.18182497 0.18340451 0.30424224 0.52475643] [0.43194502 0.29122914 0.61185289 0.13949386 0.29214465 0.36636184]]Input Embeddings (shape: (4, 6)):[[0.37454012 0.95071431 0.73199394 0.59865848 0.15601864 0.15599452] [0.05808361 0.86617615 0.60111501 0.70807258 0.02058449 0.96990985] [0.83244264 0.21233911 0.18182497 0.18340451 0.30424224 0.52475643] [0.43194502 0.29122914 0.61185289 0.13949386 0.29214465 0.36636184]]
Positional Encodings (shape: (4, 6)):[[ 0. 1. 0. 1. 0. 1. ] [ 0.84147098 0.54030231 0.2171651 0.97613297 0.04642049 0.99892228] [ 0.90929743 -0.41614684 0.42481076 0.9052784 0.09259899 0.99569976] [ 0.14112001 -0.9899925 0.6137753 0.789471 0.13830076 0.99039312]]
Input Embeddings with Positional Encoding (shape: (4, 6)):[[ 0.37454012 1.95071431 0.73199394 1.59865848 0.15601864 1.15599452] [ 0.89955459 1.40647845 0.81828011 1.68420555 0.06699498 1.96883213] [ 1.74174007 -0.20380773 0.60663573 1.08868291 0.39684123 1.52045619] [ 0.57306503 -0.69876336 1.22562819 0.92896486 0.43044541 1.35675496]]备选方案:一些 Transformer 实现使用可学习的位置嵌入(learned positional embeddings),而非固定的正弦位置嵌入。这些是训练过程中学习到的参数,类似于标记嵌入。这两种方法各有优缺点。
位置编码是 Transformer 架构的一个重要组成部分,它通过提供标记顺序信息来实现序列处理。正弦方法是一种常见且有效的方式,为每个位置提供唯一的编码,并且可以泛化到不同长度的序列。
理解位置编码对于掌握 Transformer 如何有效地建模序列数据以完成复杂的自然语言处理(NLP)任务(如翻译、摘要和生成)至关重要。它是将序列感知集成到并行处理架构中的巧妙解决方案。