TensorFlow - 词嵌入
TensorFlow - 使用 Keras 进行词嵌入
Section titled “TensorFlow - 使用 Keras 进行词嵌入”词嵌入(Word embedding)是自然语言处理(NLP)中的一种技术,它将词汇表中的单词或短语映射到实数向量。这些密集向量表示捕捉了词汇之间的语义关系,使得意思相似的词在向量空间中距离更近。这对于为机器学习模型提供有意义的输入至关重要。
例如,训练完成后,词向量可能看起来像这样(示意性的,实际值取决于训练数据和模型):
blue: (0.013, 0.001, 0.246, ..., -0.252, 1.005, 0.063)
blues: (0.014, 0.119, -0.490, ..., 0.033, -0.100, 0.116)
orange: (-0.248, -0.124, 0.210, ..., 0.080, 0.239, -0.014)
oranges: (-0.356, 0.219, 0.081, ..., -0.354, 0.385, -0.071)
请注意复数形式可能与其单数形式接近,颜色可能聚集在一起。
Word2Vec (Skip-gram 模型与负采样)
Section titled “Word2Vec (Skip-gram 模型与负采样)”Word2Vec 是一种流行的模型,用于从大型文本语料库中以无监督方式学习词嵌入。其架构之一是 Skip-gram 模型,它学习在给定目标词的情况下预测上下文词(周围的词)。为了提高训练效率,通常使用负采样(Negative Sampling),训练模型以区分真实的上下文词与少数随机选择的“负面”(不正确)词。
TensorFlow,尤其是其 Keras API,提供了实现 Word2Vec 的工具。下面是一个使用 tf.keras.layers.Embedding 和针对 NCE loss 的自定义训练循环来演示 Skip-gram 模型与负采样的示例。注意:tf.nn.nce_loss 是一个较低层的 API。对于更简单的应用,可以使用预训练词嵌入或更简单的 Keras 模型。
import osimport mathimport numpy as npimport tensorflow as tffrom tensorflow import kerasfrom tensorflow.keras import layersimport collectionsimport randomfrom tensorflow.keras.preprocessing.sequence import skipgrams
# Parametersbatch_size = 128embedding_dimension = 128 # 增加以获得更好的表示negative_samples = 5 # NCE 损失的负样本数量window_size = 2 # Skip-grams 的上下文窗口大小num_epochs = 5LOG_DIR = "logs/word2vec_keras"
if not os.path.exists(LOG_DIR): os.makedirs(LOG_DIR)
# Sample sentences (a larger, more diverse corpus is needed for meaningful embeddings)digit_to_word_map = { 1: "One", 2: "Two", 3: "Three", 4: "Four", 5: "Five", 6: "Six", 7: "Seven", 8: "Eight", 9: "Nine", 0: "Zero"}sentences = []for i in range(10000): rand_odd_ints = np.random.choice(list(range(1, 10, 2)) + [0], 3) sentences.append(" ".join([digit_to_word_map[r].lower() for r in rand_odd_ints])) rand_even_ints = np.random.choice(list(range(2, 10, 2)) + [0], 3) sentences.append(" ".join([digit_to_word_map[r].lower() for r in rand_even_ints]))
# Tokenize sentences and build vocabularywords = [word for sent in sentences for word in sent.split()]word_counts = collections.Counter(words)vocabulary_size = len(word_counts)word2index = {word: i for i, (word, _) in enumerate(word_counts.most_common())}index2word = {i: word for word, i in word2index.items()}
# Generate skip-gram pairsencoded_sentences = [[word2index[word] for word in sent.split()] for sent in sentences]
skip_gram_pairs = []for sent_indices in encoded_sentences: pairs, _ = skipgrams(sequence=sent_indices, vocabulary_size=vocabulary_size, window_size=window_size, negative_samples=0, # 我们将在损失函数中处理负采样 shuffle=True) for pair in pairs: skip_gram_pairs.append(pair)
skip_gram_pairs = np.array(skip_gram_pairs)target_words, context_words = skip_gram_pairs[:, 0], skip_gram_pairs[:, 1]
# Define the Word2Vec model using Keras Functional APIinput_target = keras.Input((1,), name='target_word')input_context = keras.Input((1,), name='context_word')
embedding_layer = layers.Embedding(vocabulary_size, embedding_dimension, name='word_embedding')
target_embedding = embedding_layer(input_target)target_embedding = layers.Reshape((embedding_dimension,))(target_embedding)
context_embedding = embedding_layer(input_context)context_embedding = layers.Reshape((embedding_dimension,))(context_embedding)
# Using dot product to calculate similarity for positive pairspositive_similarity = layers.Dot(axes=1, normalize=False)([target_embedding, context_embedding])
# For negative sampling, we'd typically use nce_loss or sample_softmax_loss in a custom training loop,# or approximate it with binary classification (true pair vs. false pair).# The original example used tf.nn.nce_loss.# A full Keras model with NCE is complex to integrate cleanly.# Typically use nce_loss or sample_softmax_loss in a custom training loop, or approximate it with binary classification (true pair vs. false pair).# 原始示例使用了 tf.nn.nce_loss。# 一个包含 NCE 的完整 Keras 模型很难干净地集成。
# Simplified approach: define a model that can be trained with binary cross-entropy# by generating explicit negative samples and labels (1 for positive, 0 for negative).# This is not exactly NCE but a common alternative.# 简化方法:定义一个可以通过生成显式负样本和标签(1 表示正样本,0 表示负样本)来使用二元交叉熵进行训练的模型。# 这并非完全是 NCE,但是一种常见的替代方案。
# For this example, we stick to a structure that *could* use tf.nn.nce_loss in a custom loop.# However, demonstrating the full custom loop with NCE is extensive for this format.# Let's focus on the embedding layer and how to extract embeddings.# 对于这个示例,我们坚持使用一个可以在自定义循环中*可能*使用 tf.nn.nce_loss 的结构。# 然而,在此格式下展示完整的带有 NCE 的自定义循环篇幅过长。# 让我们专注于嵌入层以及如何提取嵌入。
# We'll define a model that just has the embedding layer for visualization purposes,# as full NCE training is verbose.# 为了可视化目的,我们将定义一个只包含嵌入层的模型,因为完整的 NCE 训练过程比较冗长。class Word2VecModel(keras.Model): def __init__(self, vocab_size, embed_dim): super(Word2VecModel, self).__init__() self.target_embedding = layers.Embedding(vocab_size, embed_dim, input_length=1, name="w2v_embedding") # NCE loss typically requires separate weights for context predictions # NCE 损失通常需要独立的权重来进行上下文预测 self.nce_weights = layers.Dense(embed_dim, use_bias=False) # This structure is simplified. True NCE involves sampling. # 这个结构是简化的。真正的 NCE 涉及采样。
def call(self, pair): target, context = pair word_embed = self.target_embedding(target) # context_embed = self.target_embedding(context) # if using shared weights # context_embed = self.target_embedding(context) # 如果使用共享权重 # In a full NCE setup, one would compute loss against sampled negative context words. # 在完整的 NCE 设置中,需要针对采样的负面上下文词计算损失。 return word_embed # Simplified for now # 暂时简化
# Instantiate a simple embedding model for demonstration# 实例化一个简单的嵌入模型用于演示simple_embedding_model = keras.Sequential([ layers.InputLayer(input_shape=(1,)), layers.Embedding(vocabulary_size, embedding_dimension, name='embedding_layer')])
# A full training example with tf.nn.nce_loss requires a custom training loop.# Let's illustrate setting up the embedding matrix and NCE loss function.# 一个使用 tf.nn.nce_loss 的完整训练示例需要自定义训练循环。# 让我们演示如何设置嵌入矩阵和 NCE 损失函数。
# Embedding matrix (target embeddings)# 嵌入矩阵(目标词嵌入)embeddings_matrix = tf.Variable( tf.random.uniform([vocabulary_size, embedding_dimension], -1.0, 1.0), name='target_embeddings')
# NCE weights and biases (context/output embeddings)# NCE 权重和偏置(上下文/输出词嵌入)nce_weights = tf.Variable( tf.random.truncated_normal([vocabulary_size, embedding_dimension], stddev=1.0 / math.sqrt(embedding_dimension)))nce_biases = tf.Variable(tf.zeros([vocabulary_size]))
optimizer = keras.optimizers.Adam(learning_rate=0.01)
@tf.functiondef train_step(target_input, context_labels): with tf.GradientTape() as tape: # Look up embeddings for the target words # 查找目标词的嵌入 embed = tf.nn.embedding_lookup(embeddings_matrix, target_input)
# Compute NCE loss # 计算 NCE 损失 loss = tf.reduce_mean( tf.nn.nce_loss(weights=nce_weights, biases=nce_biases, labels=tf.expand_dims(context_labels, axis=1), inputs=embed, num_sampled=negative_samples, num_classes=vocabulary_size)) gradients = tape.gradient(loss, [embeddings_matrix, nce_weights, nce_biases]) optimizer.apply_gradients(zip(gradients, [embeddings_matrix, nce_weights, nce_biases])) return loss
print("Starting training with custom NCE loop...")num_steps = len(target_words) // batch_sizefor epoch in range(num_epochs): epoch_loss = 0 # Shuffle data each epoch # 每个 epoch 打乱数据 indices = np.arange(len(target_words)) np.random.shuffle(indices) shuffled_targets = target_words[indices] shuffled_contexts = context_words[indices]
for step in range(num_steps): start = step * batch_size end = (step + 1) * batch_size batch_targets = shuffled_targets[start:end] batch_contexts = shuffled_contexts[start:end]
current_loss = train_step(batch_targets, batch_contexts) epoch_loss += current_loss
if step % 500 == 0: print(f"Epoch {epoch+1}, Step {step}, Loss: {current_loss.numpy():.4f}") print(f"Epoch {epoch+1} average loss: {epoch_loss.numpy()/num_steps:.4f}")
print("Training finished.")
# Normalize embeddings before using (optional, but common for cosine similarity)# 使用前对嵌入进行归一化(可选,但对于余弦相似度很常见)final_embeddings = embeddings_matrix.numpy()norm = np.sqrt(np.sum(np.square(final_embeddings), 1, keepdims=True))normalized_embeddings = final_embeddings / norm
# Example: Find words similar to 'one'# 示例:查找与 'one' 相似的词ref_word_str = "one"if ref_word_str in word2index: ref_word_idx = word2index[ref_word_str] ref_word_embedding = normalized_embeddings[ref_word_idx]
# Calculate cosine similarity # 计算余弦相似度 similarities = np.dot(normalized_embeddings, ref_word_embedding)
# Get top N similar words # 获取最相似的 N 个词 sorted_indices = np.argsort(similarities)[::-1] print(f"\nWords similar to '{ref_word_str}':") for i in range(1, 6): # Skip the first one (itself) # 跳过第一个(自身) similar_word_idx = sorted_indices[i] similar_word_str = index2word[similar_word_idx] print(f"- {similar_word_str} (similarity: {similarities[similar_word_idx]:.3f})")else: print(f"Word '{ref_word_str}' not in vocabulary.")
# For TensorBoard embedding visualization (this part requires more setup for TF2)# You'd typically save the embeddings and metadata.tsv file.# E.g., save metadata:# 用于 TensorBoard 嵌入可视化(这部分在 TF2 中需要更多设置)# 通常需要保存嵌入和 metadata.tsv 文件。# 例如,保存元数据:with open(os.path.join(LOG_DIR, 'metadata.tsv'), "w") as f: for i in range(vocabulary_size): f.write(index2word[i] + "\n")# The embeddings themselves (normalized_embeddings) would also be saved,# and TensorBoard configured to load them. This is more involved than can be shown briefly.# See: https://www.tensorflow.org/tensorboard/tensorboard_projector_plugin# 嵌入本身(normalized_embeddings)也需要保存,# 并配置 TensorBoard 来加载它们。这比简短展示要复杂得多。# 参考:https://www.tensorflow.org/tensorboard/tensorboard_projector_plugin
print(f"\nEmbeddings and metadata.tsv (for projector) can be found in {LOG_DIR}")# To visualize, you can use the projector: http://projector.tensorflow.org/# by uploading the saved embedding tensor and metadata file.# Or configure TensorBoard with a projector_config.pbtxt.# For Keras Embedding layers, TensorBoard callback can also log embeddings.# 要可视化,可以使用 projector:http://projector.tensorflow.org/# 通过上传保存的嵌入张量和元数据文件。# 或者使用 projector_config.pbtxt 配置 TensorBoard。# 对于 Keras Embedding 层,TensorBoard 回调函数也可以记录嵌入。该脚本将:
- 从示例句子生成 Skip-gram 对。
- 定义 TensorFlow 变量用于目标嵌入、NCE 权重和偏置。
- 使用包含
tf.nn.nce_loss和 Adam 优化器的自定义循环训练这些嵌入。 - 在训练期间打印损失值。
- 训练完成后,对学习到的嵌入进行归一化。
- 最后,根据其嵌入的余弦相似度,找到并打印与 ‘one’ 最相似的词。 由于数据集小且是合成的,‘相似度’可能不具有深层语义,但应反映生成句子中的共现模式(例如,其他奇数或 ‘zero’)。使用大型真实世界文本语料库时,Word2Vec 可以学习丰富的语义关系。
现代方法通常涉及使用预训练词嵌入(例如 Word2Vec, GloVe, FastText)或训练更高级的模型,如 Transformers(例如 BERT, GPT),它们学习上下文相关的嵌入。