FIND 提示
Prompt Engineering:提取特定信息
Section titled “Prompt Engineering:提取特定信息”Prompt Engineering 中的一项常见任务是指导像 ChatGPT 这样的大语言模型(LLM)从给定文本或其自身生成的回复中提取特定的信息片段。尽管 LLM 没有字面意义上的“查找”(FIND)命令,但您可以设计 prompt 来指示它们根据您的标准定位并呈现相关细节,从而提高输出的精确性和实用性。
指导 LLM 提取信息
Section titled “指导 LLM 提取信息”您不是使用特定的指令,而是通过自然语言指示 LLM 执行提取。这包括清晰定义您正在寻找的信息,并通常指定提取数据的格式。
考虑这种概念性方法:您提供文本并要求 LLM 识别并列出其中的某些元素。
User: Here is a customer review: "I absolutely love the new QuietTune headphones! The noise cancellation is incredible, and the battery lasts for ages. My only complaint is that the carrying case feels a bit flimsy. Overall, a great purchase from AudioMax, delivered on September 15th."Please extract the product name, brand name, any positive mentions, any negative mentions, and the delivery date. Present them as a list.
LLM: Okay, here's the extracted information from the review:* **Product Name:** QuietTune headphones* **Brand Name:** AudioMax* **Positive Mentions:** Noise cancellation is incredible, battery lasts for ages, great purchase* **Negative Mentions:** Carrying case feels a bit flimsy* **Delivery Date:** September 15th在此示例中,用户的 prompt 清晰指定了要“查找”或提取的内容以及所需的输出格式。LLM 理解这些指示并相应地处理文本。
信息提取 Prompt 的最佳实践
Section titled “信息提取 Prompt 的最佳实践”为了有效指导 LLM 提取信息,请考虑以下最佳实践:
- 具体且无歧义:清晰定义您想要提取的信息。使用精确的术语。不要说‘查找重要部分’,而要说‘提取所有提到的日期’或‘识别涉及的关键人物’。
- 指定来源:指明信息是应从提供的文本中提取,还是应由 LLM 生成文本后再从其生成内容中提取。
- 定义输出格式:请求信息以结构化格式呈现,例如列表、项目符号列表、JSON 或表格。这使得输出更易于解析和使用。示例:‘提取姓名和电子邮件地址,并将其以 JSON 对象数组的形式提供。’
- 使用上下文 Prompt:在 prompt 中清晰地明确标示提供用于提取信息的文本。如果要求 LLM 先生成再提取,请确保生成指示清晰明确。
- 迭代和完善:如果初次提取不准确,请完善您的 prompt。您可能需要重新措辞、添加更多约束条件,或提供示例(少样本提示,few-shot prompting)。
- 对于复杂提取考虑思维链(Chain-of-Thought):对于多步骤提取(例如,‘查找所有提到的公司,然后为每个公司查找其 CEO’),您可以将其分解或要求 LLM‘一步一步思考’。
应用示例:用于实体提取的 Python 实现
Section titled “应用示例:用于实体提取的 Python 实现”我们来探讨一个使用 OpenAI API 从用户提供的文本中提取特定实体的 Python 脚本。
在此示例中,我们定义了一个函数 extract_entities_from_text(),该函数接受一个文本和要提取的实体类型列表。Prompt 指导 ChatGPT 在文本中查找这些实体。
from openai import OpenAIimport os
# client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))# For this example, use a placeholder. Replace with your actual key.client = OpenAI(api_key='YOUR_API_KEY')
def extract_entities_from_text(text_to_analyze, entities_to_extract): # entities_to_extract should be a string like "person names, locations, and organizations" prompt_content = f"Analyze the following text and extract all {entities_to_extract}. Present the extracted information clearly, categorizing by entity type if multiple types are requested.\n\nText to analyze:\n"""{text_to_analyze}"""\n\nExtracted Information:"
try: response = client.chat.completions.create( model="gpt-3.5-turbo", # 或者使用更新的模型 messages=[ {"role": "system", "content": "You are an AI assistant specialized in information extraction."}, {"role": "user", "content": prompt_content} ], max_tokens=300, temperature=0.2, # 降低 temperature,以获得更确定性的提取结果 n=1 ) return response.choices[0].message.content except Exception as e: return f"An error occurred: {e}"
# Example usage:user_text = "Dr. Eleanor Vance from Acme Corp visited Paris last week to meet with Jean Dupont of Beta Innovations."entities = "person names, organization names, and locations"
extracted_info = extract_entities_from_text(user_text, entities)print(extracted_info)脚本运行时,它将文本和提取指令发送给 LLM。LLM 处理后返回请求的信息。
Extracted Information:
**Person Names:*** Dr. Eleanor Vance* Jean Dupont
**Organization Names:*** Acme Corp* Beta Innovations
**Locations:*** Paris本章探讨了如何设计 prompt 来指导 LLM 提取特定信息。尽管没有字面意义上的“查找”命令,但通过清晰指定要查找的内容、源文本和所需的输出格式,您可以有效地指示 LLM 定位并呈现相关细节。
掌握这些技术需要您的指令具体明确,提供清晰的上下文,请求结构化输出,并不断迭代您的 prompt。这种能力对于从非结构化文本中进行数据挖掘、填充数据库或快速识别文档中的关键信息等任务至关重要。