Skip to content

TensorFlow - 词嵌入

词嵌入(Word embedding)是自然语言处理(NLP)中的一种技术,它将词汇表中的单词或短语映射到实数向量。这些密集向量表示捕捉了词汇之间的语义关系,使得意思相似的词在向量空间中距离更近。这对于为机器学习模型提供有意义的输入至关重要。

例如,训练完成后,词向量可能看起来像这样(示意性的,实际值取决于训练数据和模型): blue: (0.013, 0.001, 0.246, ..., -0.252, 1.005, 0.063) blues: (0.014, 0.119, -0.490, ..., 0.033, -0.100, 0.116) orange: (-0.248, -0.124, 0.210, ..., 0.080, 0.239, -0.014) oranges: (-0.356, 0.219, 0.081, ..., -0.354, 0.385, -0.071) 请注意复数形式可能与其单数形式接近,颜色可能聚集在一起。

Word2Vec 是一种流行的模型,用于从大型文本语料库中以无监督方式学习词嵌入。其架构之一是 Skip-gram 模型,它学习在给定目标词的情况下预测上下文词(周围的词)。为了提高训练效率,通常使用负采样(Negative Sampling),训练模型以区分真实的上下文词与少数随机选择的“负面”(不正确)词。

TensorFlow,尤其是其 Keras API,提供了实现 Word2Vec 的工具。下面是一个使用 tf.keras.layers.Embedding 和针对 NCE loss 的自定义训练循环来演示 Skip-gram 模型与负采样的示例。注意:tf.nn.nce_loss 是一个较低层的 API。对于更简单的应用,可以使用预训练词嵌入或更简单的 Keras 模型。

import os
import math
import numpy as np
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers
import collections
import random
from tensorflow.keras.preprocessing.sequence import skipgrams
# Parameters
batch_size = 128
embedding_dimension = 128 # 增加以获得更好的表示
negative_samples = 5 # NCE 损失的负样本数量
window_size = 2 # Skip-grams 的上下文窗口大小
num_epochs = 5
LOG_DIR = "logs/word2vec_keras"
if not os.path.exists(LOG_DIR):
os.makedirs(LOG_DIR)
# Sample sentences (a larger, more diverse corpus is needed for meaningful embeddings)
digit_to_word_map = {
1: "One", 2: "Two", 3: "Three", 4: "Four", 5: "Five",
6: "Six", 7: "Seven", 8: "Eight", 9: "Nine", 0: "Zero"
}
sentences = []
for i in range(10000):
rand_odd_ints = np.random.choice(list(range(1, 10, 2)) + [0], 3)
sentences.append(" ".join([digit_to_word_map[r].lower() for r in rand_odd_ints]))
rand_even_ints = np.random.choice(list(range(2, 10, 2)) + [0], 3)
sentences.append(" ".join([digit_to_word_map[r].lower() for r in rand_even_ints]))
# Tokenize sentences and build vocabulary
words = [word for sent in sentences for word in sent.split()]
word_counts = collections.Counter(words)
vocabulary_size = len(word_counts)
word2index = {word: i for i, (word, _) in enumerate(word_counts.most_common())}
index2word = {i: word for word, i in word2index.items()}
# Generate skip-gram pairs
encoded_sentences = [[word2index[word] for word in sent.split()] for sent in sentences]
skip_gram_pairs = []
for sent_indices in encoded_sentences:
pairs, _ = skipgrams(sequence=sent_indices,
vocabulary_size=vocabulary_size,
window_size=window_size,
negative_samples=0, # 我们将在损失函数中处理负采样
shuffle=True)
for pair in pairs:
skip_gram_pairs.append(pair)
skip_gram_pairs = np.array(skip_gram_pairs)
target_words, context_words = skip_gram_pairs[:, 0], skip_gram_pairs[:, 1]
# Define the Word2Vec model using Keras Functional API
input_target = keras.Input((1,), name='target_word')
input_context = keras.Input((1,), name='context_word')
embedding_layer = layers.Embedding(vocabulary_size,
embedding_dimension,
name='word_embedding')
target_embedding = embedding_layer(input_target)
target_embedding = layers.Reshape((embedding_dimension,))(target_embedding)
context_embedding = embedding_layer(input_context)
context_embedding = layers.Reshape((embedding_dimension,))(context_embedding)
# Using dot product to calculate similarity for positive pairs
positive_similarity = layers.Dot(axes=1, normalize=False)([target_embedding, context_embedding])
# For negative sampling, we'd typically use nce_loss or sample_softmax_loss in a custom training loop,
# or approximate it with binary classification (true pair vs. false pair).
# The original example used tf.nn.nce_loss.
# A full Keras model with NCE is complex to integrate cleanly.
# Typically use nce_loss or sample_softmax_loss in a custom training loop, or approximate it with binary classification (true pair vs. false pair).
# 原始示例使用了 tf.nn.nce_loss。
# 一个包含 NCE 的完整 Keras 模型很难干净地集成。
# Simplified approach: define a model that can be trained with binary cross-entropy
# by generating explicit negative samples and labels (1 for positive, 0 for negative).
# This is not exactly NCE but a common alternative.
# 简化方法:定义一个可以通过生成显式负样本和标签(1 表示正样本,0 表示负样本)来使用二元交叉熵进行训练的模型。
# 这并非完全是 NCE,但是一种常见的替代方案。
# For this example, we stick to a structure that *could* use tf.nn.nce_loss in a custom loop.
# However, demonstrating the full custom loop with NCE is extensive for this format.
# Let's focus on the embedding layer and how to extract embeddings.
# 对于这个示例,我们坚持使用一个可以在自定义循环中*可能*使用 tf.nn.nce_loss 的结构。
# 然而,在此格式下展示完整的带有 NCE 的自定义循环篇幅过长。
# 让我们专注于嵌入层以及如何提取嵌入。
# We'll define a model that just has the embedding layer for visualization purposes,
# as full NCE training is verbose.
# 为了可视化目的,我们将定义一个只包含嵌入层的模型,因为完整的 NCE 训练过程比较冗长。
class Word2VecModel(keras.Model):
def __init__(self, vocab_size, embed_dim):
super(Word2VecModel, self).__init__()
self.target_embedding = layers.Embedding(vocab_size, embed_dim,
input_length=1,
name="w2v_embedding")
# NCE loss typically requires separate weights for context predictions
# NCE 损失通常需要独立的权重来进行上下文预测
self.nce_weights = layers.Dense(embed_dim, use_bias=False)
# This structure is simplified. True NCE involves sampling.
# 这个结构是简化的。真正的 NCE 涉及采样。
def call(self, pair):
target, context = pair
word_embed = self.target_embedding(target)
# context_embed = self.target_embedding(context) # if using shared weights
# context_embed = self.target_embedding(context) # 如果使用共享权重
# In a full NCE setup, one would compute loss against sampled negative context words.
# 在完整的 NCE 设置中,需要针对采样的负面上下文词计算损失。
return word_embed # Simplified for now
# 暂时简化
# Instantiate a simple embedding model for demonstration
# 实例化一个简单的嵌入模型用于演示
simple_embedding_model = keras.Sequential([
layers.InputLayer(input_shape=(1,)),
layers.Embedding(vocabulary_size, embedding_dimension, name='embedding_layer')
])
# A full training example with tf.nn.nce_loss requires a custom training loop.
# Let's illustrate setting up the embedding matrix and NCE loss function.
# 一个使用 tf.nn.nce_loss 的完整训练示例需要自定义训练循环。
# 让我们演示如何设置嵌入矩阵和 NCE 损失函数。
# Embedding matrix (target embeddings)
# 嵌入矩阵(目标词嵌入)
embeddings_matrix = tf.Variable(
tf.random.uniform([vocabulary_size, embedding_dimension], -1.0, 1.0),
name='target_embeddings')
# NCE weights and biases (context/output embeddings)
# NCE 权重和偏置(上下文/输出词嵌入)
nce_weights = tf.Variable(
tf.random.truncated_normal([vocabulary_size, embedding_dimension],
stddev=1.0 / math.sqrt(embedding_dimension)))
nce_biases = tf.Variable(tf.zeros([vocabulary_size]))
optimizer = keras.optimizers.Adam(learning_rate=0.01)
@tf.function
def train_step(target_input, context_labels):
with tf.GradientTape() as tape:
# Look up embeddings for the target words
# 查找目标词的嵌入
embed = tf.nn.embedding_lookup(embeddings_matrix, target_input)
# Compute NCE loss
# 计算 NCE 损失
loss = tf.reduce_mean(
tf.nn.nce_loss(weights=nce_weights,
biases=nce_biases,
labels=tf.expand_dims(context_labels, axis=1),
inputs=embed,
num_sampled=negative_samples,
num_classes=vocabulary_size))
gradients = tape.gradient(loss, [embeddings_matrix, nce_weights, nce_biases])
optimizer.apply_gradients(zip(gradients, [embeddings_matrix, nce_weights, nce_biases]))
return loss
print("Starting training with custom NCE loop...")
num_steps = len(target_words) // batch_size
for epoch in range(num_epochs):
epoch_loss = 0
# Shuffle data each epoch
# 每个 epoch 打乱数据
indices = np.arange(len(target_words))
np.random.shuffle(indices)
shuffled_targets = target_words[indices]
shuffled_contexts = context_words[indices]
for step in range(num_steps):
start = step * batch_size
end = (step + 1) * batch_size
batch_targets = shuffled_targets[start:end]
batch_contexts = shuffled_contexts[start:end]
current_loss = train_step(batch_targets, batch_contexts)
epoch_loss += current_loss
if step % 500 == 0:
print(f"Epoch {epoch+1}, Step {step}, Loss: {current_loss.numpy():.4f}")
print(f"Epoch {epoch+1} average loss: {epoch_loss.numpy()/num_steps:.4f}")
print("Training finished.")
# Normalize embeddings before using (optional, but common for cosine similarity)
# 使用前对嵌入进行归一化(可选,但对于余弦相似度很常见)
final_embeddings = embeddings_matrix.numpy()
norm = np.sqrt(np.sum(np.square(final_embeddings), 1, keepdims=True))
normalized_embeddings = final_embeddings / norm
# Example: Find words similar to 'one'
# 示例:查找与 'one' 相似的词
ref_word_str = "one"
if ref_word_str in word2index:
ref_word_idx = word2index[ref_word_str]
ref_word_embedding = normalized_embeddings[ref_word_idx]
# Calculate cosine similarity
# 计算余弦相似度
similarities = np.dot(normalized_embeddings, ref_word_embedding)
# Get top N similar words
# 获取最相似的 N 个词
sorted_indices = np.argsort(similarities)[::-1]
print(f"\nWords similar to '{ref_word_str}':")
for i in range(1, 6): # Skip the first one (itself)
# 跳过第一个(自身)
similar_word_idx = sorted_indices[i]
similar_word_str = index2word[similar_word_idx]
print(f"- {similar_word_str} (similarity: {similarities[similar_word_idx]:.3f})")
else:
print(f"Word '{ref_word_str}' not in vocabulary.")
# For TensorBoard embedding visualization (this part requires more setup for TF2)
# You'd typically save the embeddings and metadata.tsv file.
# E.g., save metadata:
# 用于 TensorBoard 嵌入可视化(这部分在 TF2 中需要更多设置)
# 通常需要保存嵌入和 metadata.tsv 文件。
# 例如,保存元数据:
with open(os.path.join(LOG_DIR, 'metadata.tsv'), "w") as f:
for i in range(vocabulary_size):
f.write(index2word[i] + "\n")
# The embeddings themselves (normalized_embeddings) would also be saved,
# and TensorBoard configured to load them. This is more involved than can be shown briefly.
# See: https://www.tensorflow.org/tensorboard/tensorboard_projector_plugin
# 嵌入本身(normalized_embeddings)也需要保存,
# 并配置 TensorBoard 来加载它们。这比简短展示要复杂得多。
# 参考:https://www.tensorflow.org/tensorboard/tensorboard_projector_plugin
print(f"\nEmbeddings and metadata.tsv (for projector) can be found in {LOG_DIR}")
# To visualize, you can use the projector: http://projector.tensorflow.org/
# by uploading the saved embedding tensor and metadata file.
# Or configure TensorBoard with a projector_config.pbtxt.
# For Keras Embedding layers, TensorBoard callback can also log embeddings.
# 要可视化,可以使用 projector:http://projector.tensorflow.org/
# 通过上传保存的嵌入张量和元数据文件。
# 或者使用 projector_config.pbtxt 配置 TensorBoard。
# 对于 Keras Embedding 层,TensorBoard 回调函数也可以记录嵌入。

该脚本将:

  1. 从示例句子生成 Skip-gram 对。
  2. 定义 TensorFlow 变量用于目标嵌入、NCE 权重和偏置。
  3. 使用包含 tf.nn.nce_loss 和 Adam 优化器的自定义循环训练这些嵌入。
  4. 在训练期间打印损失值。
  5. 训练完成后,对学习到的嵌入进行归一化。
  6. 最后,根据其嵌入的余弦相似度,找到并打印与 ‘one’ 最相似的词。 由于数据集小且是合成的,‘相似度’可能不具有深层语义,但应反映生成句子中的共现模式(例如,其他奇数或 ‘zero’)。使用大型真实世界文本语料库时,Word2Vec 可以学习丰富的语义关系。

现代方法通常涉及使用预训练词嵌入(例如 Word2Vec, GloVe, FastText)或训练更高级的模型,如 Transformers(例如 BERT, GPT),它们学习上下文相关的嵌入。