Transformers 中的输入嵌入
Transformer 模型中的输入 Embedding
Section titled “Transformer 模型中的输入 Embedding”Transformer 架构作为现代自然语言处理(NLP)的基石,由编码器和解码器组件构成,每个组件包含多个机制和子层。在 Transformer 中处理文本输入的第一个步骤是 Input Embedding 层(输入嵌入层)。
Input Embedding(输入嵌入)是一个关键组件,它将原始文本数据(例如单词或子词单元,即 token)转换为模型可以处理的连续向量表示。这些稠密向量捕获了 token 的语义信息,使模型能够有效地理解和处理文本。
本章将解释什么是 Input Embedding,它们为何至关重要,以及如何在 Transformer 中实现它们,并提供一个概念性的 Python 示例来说明这些概念。
什么是 Input Embedding?
Section titled “什么是 Input Embedding?”Input Embedding 是离散输入 token(例如,单词、子词或字符)的向量表示。这些向量旨在捕获 token 的语义意义和上下文细微差别,使模型能够辨别它们之间的关系。例如,具有相似含义的单词在 Embedding 空间中应该具有相似的向量表示。
Input Embedding 子层在 Transformer 中的作用是将这些输入 token 映射到高维向量空间中,通常其维度表示为 $d_{model}$(例如,$d_{model} = 512$ 或 $d_{model} = 768$)。在这个空间中,具有相似语义属性的 token 位置更接近。
Input Embedding 在 Transformer 中的重要性
Section titled “Input Embedding 在 Transformer 中的重要性”让我们了解为什么 Input Embedding 在 Transformer 模型中至关重要:
语义表示 (Semantic Representation)
Section titled “语义表示 (Semantic Representation)”Input Embedding 捕获输入文本中单词或 token 之间的语义相似性。例如,像 “king” 和 “queen”,或者 “apple” 和 “orange” 这样的单词的向量在 Embedding 空间中可能相对接近,反映了它们的语义关系。
维度和效率 (Dimensionality and Efficiency)
Section titled “维度和效率 (Dimensionality and Efficiency)”传统的 one-hot 编码方法将每个 token 表示为一个稀疏的二元向量,其中只有一个维度是“热”(设置为 1)。对于大型词汇表,这会变得非常低效,导致维度极高且稀疏的表示。Input Embedding 提供了一种紧凑的稠密表示(例如,对于包含 5 万个单词的词汇表,是 512 维而不是 5 万维),从而降低了计算复杂性和内存需求。
增强学习和泛化能力 (Enhanced Learning and Generalization)
Section titled “增强学习和泛化能力 (Enhanced Learning and Generalization)”因为 Embedding 可以捕获上下文关系和语义细微差别,它们提高了模型学习模式和从训练数据泛化的能力。这使得模型在各种 NLP 任务(如翻译、文本摘要和问答)上表现更好。
Input Embedding 子层的工作原理
Section titled “Input Embedding 子层的工作原理”Transformer 中生成 Input Embedding 的过程,类似于其他序列转换模型,涉及以下步骤:
步骤 1:分词 (Tokenization)
Section titled “步骤 1:分词 (Tokenization)”在对输入 token 进行 Embedding 之前,必须对原始输入文本进行分词 (tokenization)。分词是将文本分解为更小单元(称为 token)的过程。常见的分词策略包括:
- 单词级分词 (Word-level tokenization): 将文本分割成独立的单词。这可能导致词汇表过大以及处理词汇外 (out-of-vocabulary, OOV) 词汇的问题。
- 子词级分词 (Subword-level tokenization): 将文本分割成更小的单元,可以是单词的一部分或频繁出现的字符序列。常用的算法包括 Byte Pair Encoding (BPE)、WordPiece 和 SentencePiece。这种方法能有效处理 OOV 词汇并管理词汇表大小。例如,文本 “Transformers revolutionized NLP” 可能会被分词为
["Transform", "##ers", "revolutionized", "NLP"]。
步骤 2:Embedding 层 (Embedding Layer)
Section titled “步骤 2:Embedding 层 (Embedding Layer)”Embedding 层本质上是一个查找表(或一个矩阵),它将每个 token ID(在词汇表中代表唯一 token 的整数)映射到一个固定维度的稠密向量。这个过程包括:
- 词汇表 (Vocabulary): 模型识别的一组预定义的唯一 token。每个 token 被分配一个唯一的整数 ID。
- Embedding 维度 ($d_{model}$): Token 在其中表示的向量空间的大小(例如,512)。每个 token ID 映射到一个 $d_{model}$ 维向量。
当一个 token ID 被传递给 Embedding 层时,它会从 Embedding 矩阵中检索对应的稠密向量。这个矩阵通常是随机初始化的,然后在模型训练过程中进行学习(更新)。
Input Embedding 的概念性 Python 实现
Section titled “Input Embedding 的概念性 Python 实现”下面是一个使用 NumPy 的简化 Python 示例,用于说明 Input Embedding 的概念。在实践中,像 PyTorch 或 TensorFlow 这样的深度学习框架,以及像 Hugging Face Transformers 这样的库,在更大的模型架构中更高效地处理这些操作。
import numpy as np
# Example text and basic word tokenizationtext = "Transformers revolutionized the field of NLP"tokens = text.lower().split() # Simple tokenization by splitting on space and lowercasing
# Creating a vocabulary (mapping words to unique integer IDs)vocab = {word: idx for idx, word in enumerate(set(tokens))} # Use set to get unique words
# Example input (sequence of token indices based on our vocabulary)# Ensure all tokens from the original sentence are in the vocab before mappinginput_indices = np.array([vocab[word] for word in tokens if word in vocab])
print("Vocabulary:", vocab)print("Input Indices:", input_indices)
# Parameters for the embedding layervocab_size = len(vocab)embed_dim = 4 # Using a very small dimension for illustration (e.g., d_model = 512 in real Transformers)
# Initialize the embedding matrix with random values (this would be learned during training)# Each row corresponds to the embedding of a word (token ID)embedding_matrix = np.random.rand(vocab_size, embed_dim)
# Get the embeddings for the input indices (lookup operation)input_embeddings = embedding_matrix[input_indices]
print("\nEmbedding Matrix (Shape: {}, {}):\n".format(vocab_size, embed_dim), embedding_matrix)print("\nInput Embeddings (Shape: {}, {}):\n".format(len(input_indices), embed_dim), input_embeddings)
# For further learning, explore:# - Subword tokenization (BPE, WordPiece) using libraries like Hugging Face Tokenizers.# - torch.nn.Embedding in PyTorch or tf.keras.layers.Embedding in TensorFlow.这个 Python 脚本首先对文本进行分词并创建一个词汇表,将每个唯一的单词映射到一个整数 ID。然后,它初始化一个包含随机值的 Embedding 矩阵,其中每一行代表一个单词(token ID)的 Embedding 向量。这里的 embed_dim 为了方便显示而设置得很小;实际的 Transformer 使用更大的维度(例如 512)。最后,它检索输入序列的 Embedding。
Vocabulary: {'nlp': 0, 'field': 1, 'transformers': 2, 'of': 3, 'the': 4, 'revolutionized': 5}# 注意:由于使用了 set,顺序可能不同Input Indices: [2 5 4 1 3 0] # 对应于 'transformers', 'revolutionized', 'the', 'field', 'of', 'nlp'
Embedding Matrix (Shape: 6, 4): [[0.123... 0.456... 0.789... 0.101...] [0.234... 0.567... 0.890... 0.202...] ... [0.789... 0.101... 0.303... 0.606...]] # 随机值
Input Embeddings (Shape: 6, 4): [[...对应于 'transformers' 的行...] [...对应于 'revolutionized' 的行...] ... [...对应于 'nlp' 的行...]] # 从 Embedding Matrix 中选取的行实际输出中的随机数每次运行脚本时都会有所不同。关键要点是从 token ID 到稠密向量的映射关系。
Input Embedding 子层是 Transformer 模型的基础步骤,它将原始文本数据转换为模型可以处理的有意义的数值表示。通过捕获语义相似性并提供稠密高效的向量,Input Embedding 使 Transformer 能够有效地学习文本数据中复杂的模式和关系。
理解分词策略和 Embedding 层的概念对于任何使用或构建用于 NLP 任务的 Transformer 模型的人来说至关重要。虽然本章提供了概念性概述和一个基本的 Python 示例,但实际应用涉及更复杂的 Tokenizer 和集成在深度学习框架中的 Embedding 技术。