Skip to content

使用 Python 进行 AI – NLTK 包

Natural Language Toolkit (NLTK) 是一个用于各种自然语言处理(Natural Language Processing, NLP)任务的综合性 Python 库。本章将指导您完成 NLTK 的设置并探索其核心功能。

由于人类语言固有的歧义性和上下文相关性,自然语言处理(NLP)应用程序面临挑战。NLTK 提供了工具和资源来构建能够以更接近人类的方式处理和理解文本的应用程序。由于其丰富的 datasets (corpora) 和用于分词(tokenization)、词干提取(stemming)、标注(tagging)、解析(parsing)和分类(classification)等任务的模块化组件,它对于学习 NLP 概念尤其有价值。

您可以使用 Python 的包安装器 pip 来安装 NLTK:

pip install nltk

如果您使用 Anaconda,可以通过 conda 安装:

conda install -c anaconda nltk

安装后,在您的 Python 脚本或交互式会话中导入 NLTK:

import nltk

NLTK 附带了许多语料库(corpora)、预训练模型和其他数据资源。您需要单独下载这些资源。您可以在 Python 中交互式地完成此操作:

nltk.download()

此命令将打开 NLTK 下载器窗口。您可以选择下载单个软件包(如用于分词的 ‘punkt’,用于词形还原的 ‘wordnet’,‘stopwords’)或常见的集合,如 ‘book’(NLTK 书中示例所需的所有数据)。对于一般用途,下载 ‘popular’ 或 ‘all’ 也是一种选择,尽管 ‘all’ 文件可能相当大。

虽然 NLTK 功能强大,但您也可能需要结合使用其他库或用于更高级的任务:

gensim 是一个用于主题建模(topic modeling)、文档相似度分析以及处理词嵌入(word embeddings)(如 Word2Vec)的强大库。使用以下命令安装:

pip install gensim

spaCy 是一个工业级的 NLP 库,以其速度和效率著称,提供了用于各种语言和任务(如命名实体识别 (Named Entity Recognition, NER)、词性标注 (Part-of-Speech, POS) tagging 和依存句法分析)的预训练模型。使用以下命令安装:

pip install spacy
python -m spacy download en_core_web_sm # 下载一个小型英文模型

scikit-learn 提供了用于文本特征提取(例如 CountVectorizer, TfidfVectorizer)和适用于文本分类任务的各种机器学习算法的工具。使用以下命令安装:

pip install scikit-learn

基础 NLP 任务:分词、词干提取和词形还原

Section titled “基础 NLP 任务:分词、词干提取和词形还原”

这些是许多 NLP 管道中基础的预处理步骤。

分词(tokenization)是将文本(字符序列)分解为更小单元(称为标记,tokens)的过程。这些标记通常是词语、数字或标点符号。这也称为词语切分(word segmentation)。

例如,输入字符串:“NLTK is great for learning NLP! It costs $0.”

可以分词为:['NLTK', 'is', 'great', 'for', 'learning', 'NLP', '!', 'It', 'costs', '$', '0', '.']

NLTK 提供了几种分词器:

  • 句子分词器(Sentence Tokenizer, sent_tokenize):将文本分割成句子。
  • 词语分词器(Word Tokenizer, word_tokenize):将文本(或句子)分割成词语和标点符号。
  • WordPunctTokenizer:基于空格分割文本,并将标点符号与相邻词语分开。

使用 word_tokenize 和 sent_tokenize 的示例(需要 ‘punkt’ 数据包):

from nltk.tokenize import sent_tokenize, word_tokenize
# 确保已下载 'punkt':nltk.download('punkt')
text_example = "NLTK is great for learning NLP! It costs $0. What do you think?"
sentences = sent_tokenize(text_example)
print(f"句子: {sentences}")
first_sentence_words = word_tokenize(sentences[0])
print(f"第一个句子中的词语: {first_sentence_words}")

词干提取(stemming)是一种启发式过程,通过去除派生或屈折词缀(前缀/后缀)将词语简化为其“词干”或基本形式。目标是将一个词语的不同形式归入一个共同的表示。例如,“running”、“runs”、“ran”可能都会被提取为“run”。

NLTK 提供了几种词干提取器:

  • PorterStemmer:最古老、最广泛使用的词干提取器之一。它相对温和。
  • LancasterStemmer:一个更激进的词干提取器,通常产生更短、有时不那么直观的词干。
  • SnowballStemmer:支持多种语言,通常比 Porter 更好。它在激进性和准确性之间取得了良好的平衡。

使用 PorterStemmer 的示例:

from nltk.stem.porter import PorterStemmer
porter = PorterStemmer()
words_to_stem = ["running", "studies", "leaves", "activity", "connection"]
stemmed_words = [porter.stem(word) for word in words_to_stem]
print(f"词干提取结果 (Porter): {stemmed_words}")
# 输出: 词干提取结果 (Porter): ['run', 'studi', 'leav', 'activ', 'connect']

词形还原(lemmatization)旨在将词语还原到其字典形式,称为“词元”(lemma)。与词干提取不同,词形还原考虑词语的含义和词性。它使用词汇表(如 WordNet)和形态分析。

例如,“better”会被还原为“good”(如果其词性是形容词),“is”、“are”、“was”会被还原为“be”。“leaves”可以被还原为“leaf”(名词)或“leave”(动词)。

NLTK 的 WordNetLemmatizer 是常用的(需要 ‘wordnet’ 数据包):

from nltk.stem import WordNetLemmatizer
from nltk.corpus import wordnet # 用于指定词形还原器的词性
# 确保已下载 'wordnet' 和 'omw-1.4'(Open Multilingual Wordnet):
# nltk.download('wordnet')
# nltk.download('omw-1.4')
lemmatizer = WordNetLemmatizer()
print(f"对 'studies' 进行词形还原: {lemmatizer.lemmatize('studies', pos=wordnet.NOUN)}") # 'study'
print(f"对 'studies' 进行词形还原(动词): {lemmatizer.lemmatize('studies', pos=wordnet.VERB)}") # 'study'
print(f"对 'better' 进行词形还原: {lemmatizer.lemmatize('better', pos=wordnet.ADJ)}") # 'good'
print(f"对 'is' 进行词形还原: {lemmatizer.lemmatize('is', pos=wordnet.VERB)}") # 'be'

词形还原通常比词干提取产生更符合语言学规则的基本形式,但计算成本更高。

分块(chunking)或浅层解析(shallow parsing)是将文本识别并分割成语法相关的词语组的过程,例如名词短语(Noun Phrases, NP)、动词短语(Verb Phrases, VP)等。它不构建完整的解析树,但提供了一种更简单、通常更鲁棒的结构分析。

例如,在“The clever fox was jumping over the lazy dog.”这句话中,分块可能会将“The clever fox”和“the lazy dog”识别为名词短语。

要执行分块,您通常首先需要获取标记的词性(Part-of-Speech, POS)标签。NLTK 提供了词性标注器(POS taggers)(例如,nltk.pos_tag,需要 ‘averaged_perceptron_tagger’)。

步骤:

  1. 对句子进行分词。
  2. 对标记进行词性标注。
  3. 定义分块语法(通常使用基于词性标签的正则表达式)。
  4. 使用此语法创建分块解析器。
  5. 将解析器应用于已标注的标记。
import nltk
# 确保已下载所需的 NLTK 数据:
# nltk.download('averaged_perceptron_tagger')
# nltk.download('punkt')
sentence_text = "The clever fox was jumping over the lazy dog."
tokens = nltk.word_tokenize(sentence_text)
tagged_tokens = nltk.pos_tag(tokens)
print(f"已标注标记: {tagged_tokens}")
# 定义名词短语(NP)的语法
# NP:一个可选的限定词(DT),任意数量的形容词(JJ*),然后是一个名词(NN)
# 或者一个专有名词(NNP),或者代词(PRP)
grammar = r"""
NP: {<DT>?<JJ>*<NN>}
{<DT>?<JJ>*<NNS>}
{<DT>?<JJ>*<NNP>}
{<DT>?<JJ>*<NNPS>}
{<PRP>}
"""
chunk_parser = nltk.RegexpParser(grammar)
chunked_sentence_tree = chunk_parser.parse(tagged_tokens)
print("\n分块句子树:")
# 以文本形式(而不是图形形式)打印树结构:
chunked_sentence_tree.pprint() # 美观打印
# 只显示分块:
# for subtree in chunked_sentence_tree.subtrees():
# if subtree.label() == 'NP':
# print(subtree)

chunked_sentence_tree.pprint() 的输出看起来会像这样,显示了树结构和识别出的名词短语:

(S
(NP The/DT clever/JJ fox/NN)
was/VBD
jumping/VBG
over/IN
(NP the/DT lazy/JJ dog/NN)
./.)

这种文本表示显示,“The clever fox”和“the lazy dog”已被正确识别为名词短语(NP)。

词袋(Bag-of-Words, BoW)模型是一种从文本中提取特征的基本技术。它将文本(如句子或文档)表示为其词语的无序集合(一个“袋子”),忽略语法和词序,但记录词语频率。

  • 词汇表创建: 创建语料库(文档集合)中所有唯一词语的词汇表(vocabulary)。
  • 向量化: 每个文档被表示为一个数值向量。该向量的长度是词汇表的大小。每个维度对应词汇表中的一个词语,该维度的值通常是该词语在文档中出现的次数。

示例:

  • 句子 1:“The cat sat on the mat.”
  • 句子 2:“The dog chased the cat。”

词汇表:['The', 'cat', 'sat', 'on', 'mat', 'dog', 'chased'] (为简单起见,忽略大小写和标点符号)

向量:

  • 句子 1:[2, 1, 1, 1, 1, 0, 0] (分别对应 ‘The’, ‘cat’, ‘sat’, ‘on’, ‘mat’, ‘dog’, ‘chased’)
  • 句子 2:[1, 1, 0, 0, 0, 1, 1]

词袋模型使用原始计数。TF-IDF 是一种改进,它反映了词语在文档集合(corpus)中的重要程度。

衡量词语在文档中出现的频率。TF(t, d) = (词语 t 在文档 d 中出现的次数) / (文档 d 中总词语数)。

衡量词语的重要性。它会降低在许多文档中频繁出现的词语(如“the”、“is”)的权重,并提高稀有词语的权重。IDF(t, D) = log(文档总数 D / 包含词语 t 的文档数)。

词语 t 在文档 d 中的 TF-IDF 分数是 TF(t,d) * IDF(t,D)。

使用 scikit-learn 构建 BoW 和 TF-IDF 模型

Section titled “使用 scikit-learn 构建 BoW 和 TF-IDF 模型”

scikit-learn 提供了方便的工具来构建 BoW 和 TF-IDF 模型:

from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
sentences = [
'This is the first document.',
'This document is the second document.',
'And this is the third one.',
'Is this the first document?'
]
# 词袋模型 (CountVectorizer)
count_vectorizer = CountVectorizer()
bow_matrix = count_vectorizer.fit_transform(sentences)
print("\nBoW 词汇表:", count_vectorizer.vocabulary_)
print("BoW 矩阵 (稀疏):")
print(bow_matrix.toarray())
# TF-IDF (TfidfVectorizer)
tfidf_vectorizer = TfidfVectorizer()
tfidf_matrix = tfidf_vectorizer.fit_transform(sentences)
print("\nTF-IDF 词汇表:", tfidf_vectorizer.vocabulary_)
print("TF-IDF 矩阵 (稀疏):")
print(tfidf_matrix.toarray())

这些矩阵(BoW 或 TF-IDF)作为特征向量,用于训练机器学习模型。

NLTK & scikit-learn 的应用与问题解决

Section titled “NLTK & scikit-learn 的应用与问题解决”

1. 文本类别预测(新闻组分类)

Section titled “1. 文本类别预测(新闻组分类)”

预测文档的类别(例如,体育、政治、科技)。我们将使用来自 scikit-learn 的 20 Newsgroups 数据集,结合 TF-IDF 和朴素贝叶斯(Naive Bayes)分类器。

from sklearn.datasets import fetch_20newsgroups
from sklearn.naive_bayes import MultinomialNB
from sklearn.pipeline import make_pipeline # 用于链接向量化器和分类器
# 定义要加载的类别以便简化
categories = ['alt.atheism', 'soc.religion.christian', 'comp.graphics', 'sci.med']
# 获取训练数据
train_data = fetch_20newsgroups(subset='train', categories=categories, shuffle=True, random_state=42)
# 创建一个管道:TF-IDF 向量化器 -> 多项式朴素贝叶斯分类器
model = make_pipeline(TfidfVectorizer(), MultinomialNB())
# 训练模型
model.fit(train_data.data, train_data.target)
# 预测新文本示例
new_texts = [
"God is love",
"OpenGL on the GPU is fast",
"Christianity and science can coexist",
"Medical imaging with advanced algorithms"
]
predicted_categories_indices = model.predict(new_texts)
print("\n类别预测:")
for text, index in zip(new_texts, predicted_categories_indices):
print(f'输入: "{text}" => 预测类别: {train_data.target_names[index]}')

此示例演示了一个基本的文本分类管道。

2. 从姓名识别性别(示例性示例)

Section titled “2. 从姓名识别性别(示例性示例)”

一个基于姓名预测性别的简单分类器,使用姓名最后一个字母等特征。NLTK 的 names 语料库提供了带标签的数据。

import random
from nltk.corpus import names # nltk.download('names')
from nltk import NaiveBayesClassifier
from nltk.classify import accuracy as nltk_accuracy
# 特征提取器:使用姓名的最后 N 个字母
def gender_features(word, n=2):
return {'last_letters': word[-n:].lower()}
# 准备数据
# 确保已下载 names 语料库
male_names = [(name, 'male') for name in names.words('male.txt')]
female_names = [(name, 'female') for name in names.words('female.txt')]
all_names = male_names + female_names
random.shuffle(all_names)
# 创建特征集
N_letters = 2 # 使用最后 2 个字母作为特征
featuresets = [(gender_features(name, N_letters), gender) for (name, gender) in all_names]
# 分割为训练集和测试集
train_size = int(0.8 * len(featuresets))
train_set, test_set = featuresets[:train_size], featuresets[train_size:]
# 训练朴素贝叶斯分类器
classifier = NaiveBayesClassifier.train(train_set)
# 测试准确率
print(f"\n性别识别器(特征:最后 {N_letters} 个字母):")
print(f"准确率: {nltk_accuracy(classifier, test_set) * 100:.2f}%")
# 对新姓名进行预测
names_to_test = ['Neo', 'Trinity', 'Alex', 'Jordan', 'Taylor']
for name in names_to_test:
guess = classifier.classify(gender_features(name, N_letters))
print(f"姓名: {name}, 预测性别: {guess}")

这是一个简化的模型,其准确率将有限。可靠的性别识别(本身是一个复杂且细致的任务)需要更复杂的特征和模型。

主题建模(topic modeling)是一种无监督技术,用于发现文档集合中抽象的“主题”或隐藏的主题结构。每个主题是一个词语分布,而每个文档是主题的混合。

  • 潜在狄利克雷分配(Latent Dirichlet Allocation, LDA):一种概率生成模型。最常用。gensim 提供了良好的实现。
  • 潜在语义分析(Latent Semantic Analysis, LSA/LSI):在文档-词项矩阵上使用矩阵分解(SVD)。
  • 非负矩阵分解(Non-Negative Matrix Factorization, NMF):另一种矩阵分解技术。
  • 文档组织与浏览:按主题对文档进行分组。
  • 改进文本分类:使用主题作为特征。
  • 推荐系统:根据共享主题推荐内容。

使用 gensim 进行 LDA 的示例(概念性大纲 - 需要数据预处理):

from gensim.corpora import Dictionary
from gensim.models import LdaModel
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords # nltk.download('stopwords')
import string
# 示例文档(替换为您的实际文本数据)
documents_sample = [
"Sugar is bad to consume. My sister likes to have sugar, but not my father.",
"My father spends a lot of time driving my sister around to dance practice.",
"Doctors suggest that driving may cause increased stress and blood pressure.",
"Sometimes I feel pressure to perform well at school, but my father never seems to drive my sister to do better.",
"Health experts say that Sugar is not good for your lifestyle."
]
# 预处理(基本示例)
# nltk.download('stopwords')
# nltk.download('punkt')
stop_words = set(stopwords.words('english'))
punctuation = set(string.punctuation)
processed_docs = []
for doc in documents_sample:
tokens = word_tokenize(doc.lower())
cleaned_tokens = [token for token in tokens if token.isalpha() and token not in stop_words and token not in punctuation]
processed_docs.append(cleaned_tokens)
# 为 LDA 创建字典和语料库
id2word = Dictionary(processed_docs)
corpus = [id2word.doc2bow(doc) for doc in processed_docs]
# 构建 LDA 模型(例如,查找 2 个主题)
if corpus and id2word: # 确保语料库和字典不为空
lda_model = LdaModel(corpus=corpus, id2word=id2word, num_topics=2, random_state=100,
update_every=1, chunksize=100, passes=10, alpha='auto', per_word_topics=True)
print("\nLDA 主题(示例):")
for idx, topic in lda_model.print_topics(-1):
print(f"主题: {idx} \n词语: {topic}")
else:
print("\nLDA 主题:语料库或字典为空,无法训练 LDA 模型。")

此示例展示了基本步骤:预处理文本,创建字典(将词语映射到 ID)和 BoW 语料库(corpus),然后训练 LDA 模型。输出将显示与每个发现的主题关联的顶部词语。实际的主题建模涉及更细致的预处理和超参数调整。