Requests - 使用 Requests 进行网页抓取
Requests - 使用 Requests 和 BeautifulSoup 进行网页抓取
Section titled “Requests - 使用 Requests 和 BeautifulSoup 进行网页抓取”网页抓取(Web scraping)是自动从网站提取数据的过程。requests 库是一个强大的工具,用于获取网页的原始 HTML 内容。然而,要解析此 HTML(即导航其结构并提取特定信息),您通常会使用专门的解析库。BeautifulSoup 是一个非常流行的 Python 库,用于执行此任务,它可以与 requests 很好地协同工作。
网页抓取的道德考量
Section titled “网页抓取的道德考量”在开始任何网页抓取项目之前,理解并尊重道德和法律界限至关重要:
- 检查
robots.txt:大多数网站都提供一个robots.txt文件(例如,domain.com/robots.txt),该文件指定了针对网络爬虫和机器人的规则。始终尊重这些规则。 - 速率限制:不要在短时间内向服务器发送过多请求。这可能导致服务器过载,并可能导致您的 IP 被封锁。在请求之间实施延迟(例如,
time.sleep())。 - 标识您的爬虫:在请求中设置描述性的
User-Agent头部。这有助于网站管理员识别您的抓取程序,并在必要时与您联系(例如,User-Agent: MyCoolScraper/1.0 (+http://mywebsite.com/scraper_info))。 - 查阅服务条款:许多网站在其服务条款中明确说明了数据抓取政策。确保您的抓取活动符合规定。
- 数据隐私和版权:注意个人数据和版权限制。仅抓取公开可用的数据并负责任地使用它们。
- 轻量抓取:仅请求您需要的数据,并考虑是否有官方 API 可用,这始终是抓取的更好替代方案。
负责任的抓取是维持访问和良好网络公民身份的关键。
设置您的环境
Section titled “设置您的环境”要开始抓取,您需要安装 requests 和 beautifulsoup4:
pip install requests beautifulsoup4BeautifulSoup 可以使用不同的 HTML 解析器。Python 内置的 html.parser 是一个不错的默认选项。为了可能更快的解析速度和对格式错误的 HTML 更高的容错性,您可以安装 lxml:
pip install lxml如果安装了 lxml,在创建 Soup 对象时指定 'lxml',BeautifulSoup 可以自动使用它。
基本的抓取流程
Section titled “基本的抓取流程”典型的网页抓取过程包括以下步骤:
- 使用
requests.get(url, headers=...)下载目标网页的 HTML 内容。 - 检查
response.status_code以确保请求成功(例如,200 OK)。使用response.raise_for_status()进行简单的错误检查。 - 从
response.text创建一个BeautifulSoup对象,指定解析器(例如,'html.parser'或'lxml')。 - 使用 BeautifulSoup 的方法(如
find()、find_all()、CSS 选择器)在解析后的 HTML 树中定位和提取所需的数据元素。
示例:抓取 ‘quotes.toscrape.com’ 的引言
Section titled “示例:抓取 ‘quotes.toscrape.com’ 的引言”quotes.toscrape.com 是一个专为练习网页抓取而设计的网站。我们将编写一个脚本来提取引言、其作者和标签。
import requestsfrom bs4 import BeautifulSoupimport time # For respectful scraping# 为了礼貌地进行抓取
URL = 'http://quotes.toscrape.com/'# Define a user-agent to be polite# 定义一个 User-Agent,以示友好HEADERS = { 'User-Agent': 'MyTutorialScraper/1.0 (contact@example.com; +http://example.com/about_scraper)'}
try: print(f'Fetching URL: {URL}') print(f'正在抓取 URL: {URL}') response = requests.get(URL, headers=HEADERS, timeout=10) # Always use a timeout # 始终使用超时设置 response.raise_for_status() # Raise an exception for bad status codes (4xx or 5xx) # 对于不成功的状态码(4xx 或 5xx)引发异常
# Parse the HTML content using BeautifulSoup # 使用 BeautifulSoup 解析 HTML 内容 # Using 'lxml' if available, otherwise 'html.parser' # 如果 lxml 可用则使用它,否则使用 html.parser try: soup = BeautifulSoup(response.text, 'lxml') except ImportError: soup = BeautifulSoup(response.text, 'html.parser')
# Extract the page title (as an example) # 提取页面标题(作为示例) page_title = soup.title.string if soup.title else 'No title found' page_title = soup.title.string if soup.title else '未找到标题' print(f'Page Title: {page_title}\n') print(f'页面标题: {page_title}\n')
# Find all quote containers (each quote is in a <div> with class='quote') # 查找所有引言容器(每个引言都在一个 class='quote' 的 <div> 中) quote_elements = soup.find_all('div', class_='quote')
if not quote_elements: print('No quotes found. The website structure might have changed or content is not as expected.') print('未找到引言。网站结构可能已更改或内容与预期不符。') else: print(f'Found {len(quote_elements)} quotes on the page. Displaying first few:') print(f'页面上找到 {len(quote_elements)} 条引言。显示前几条:') for i, quote_div in enumerate(quote_elements[:3]): # Display first 3 quotes for brevity # 为简洁起见,显示前 3 条引言 text_element = quote_div.find('span', class_='text') author_element = quote_div.find('small', class_='author') tags_container = quote_div.find('div', class_='tags')
quote_text = text_element.get_text(strip=True) if text_element else 'N/A' quote_author = author_element.get_text(strip=True) if author_element else 'N/A'
tags_list = [] if tags_container: tag_elements = tags_container.find_all('a', class_='tag') tags_list = [tag.get_text(strip=True) for tag in tag_elements]
print(f'\nQuote {i+1}:') print(f'\n引言 {i+1}:') print(f' Text: "{quote_text}"') print(f' Author: {quote_author}') print(f' Tags: {", ".join(tags_list) if tags_list else "N/A"}') print(f' 文本: “{quote_text}”') print(f' 作者: {quote_author}') print(f' 标签: {", ".join(tags_list) if tags_list else "N/A"}')
# Respectful delay if scraping multiple pages/items # time.sleep(1) # 如果抓取多个页面/项目,请设置礼貌的延迟
except requests.exceptions.HTTPError as http_err: print(f'HTTP error occurred while fetching {URL}: {http_err}') print(f'抓取 {URL} 时发生 HTTP 错误: {http_err}')except requests.exceptions.ConnectionError as conn_err: print(f'Connection error occurred while fetching {URL}: {conn_err}') print(f'抓取 {URL} 时发生连接错误: {conn_err}')except requests.exceptions.Timeout as timeout_err: print(f'Timeout error occurred while fetching {URL}: {timeout_err}') print(f'抓取 {URL} 时发生超时错误: {timeout_err}')except requests.exceptions.RequestException as req_err: print(f'An general error occurred during the request to {URL}: {req_err}') print(f'请求 {URL} 期间发生一般错误: {req_err}')except Exception as e: print(f'An unexpected error occurred during scraping: {e}') print(f'抓取期间发生意外错误: {e}')预期输出(根据当前网站内容和显示的引言数量可能略有不同):
Section titled “预期输出(根据当前网站内容和显示的引言数量可能略有不同):”Fetching URL: http://quotes.toscrape.com/Page Title: Quotes to Scrape
Found 10 quotes on the page. Displaying first few:
Quote 1: Text: "The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking." Author: Albert Einstein Tags: change, deep-thoughts, thinking, world
Quote 2: Text: "It is our choices, Harry, that show what we truly are, far more than our abilities." Author: J.K. Rowling Tags: abilities, choices
Quote 3: Text: "There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle." Author: Albert Einstein Tags: inspirational, life, live, miracle, miracles(Actual quotes and tags on the site may change.)
正在抓取 URL: http://quotes.toscrape.com/页面标题: Quotes to Scrape
页面上找到 10 条引言。显示前几条:
引言 1: 文本: “The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.” 作者: Albert Einstein 标签: change, deep-thoughts, thinking, world
引言 2: 文本: “It is our choices, Harry, that show what we truly are, far more than our abilities.” 作者: J.K. Rowling 标签: abilities, choices
引言 3: 文本: “There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.” 作者: Albert Einstein 标签: inspirational, life, live, miracle, miracles(网站上的实际引言和标签可能会发生变化。)常用的 BeautifulSoup 数据提取操作
Section titled “常用的 BeautifulSoup 数据提取操作”soup.find(‘tag_name’, attrs={‘attribute_name’: ‘value’}, class_=‘css_class_name’): 查找第一个匹配指定条件的元素。soup.find_all(‘tag_name’, class_=‘css_class_name’, limit=N): 查找所有匹配的元素(如果指定了limit,则最多查找 N 个),并将它们作为列表 (ResultSet) 返回。element.get_text(strip=True, separator=’ ’): 从元素及其子元素中提取所有人类可读的文本。strip=True删除每段文本的首尾空格。separator可以用于连接文本片段。element[‘attribute_name’](例如,link_tag['href']):访问元素的属性值。- CSS 选择器:
soup.select(‘div.content > p’)允许使用 CSS 选择器语法查找元素,这非常强大。例如,soup.select_one('#main-id .item-class')查找匹配此 CSS 选择器的第一个元素。 - 遍历树:
.parent,.children(迭代器),.find_next_sibling(),.find_previous_sibling()等,允许遍历解析后的 HTML 结构。
有关 BeautifulSoup 功能的全面指南,请参阅其官方文档。
挑战:动态内容和反爬虫措施
Section titled “挑战:动态内容和反爬虫措施”许多现代网站在初始页面加载后使用 JavaScript 动态加载内容。由于 requests(以及由 requests 提供数据的 BeautifulSoup)只抓取初始 HTML 源代码,并且不执行 JavaScript,您可能会发现所需的数据缺失。
处理动态内容的策略包括:
- **检查网络请求:**使用浏览器的开发者工具(Network 选项卡)查看 JavaScript 是否向 API 端点发起了 XHR/Fetch 请求,这些请求返回数据(通常是 JSON 格式)。如果是这样,您通常可以使用
requests直接调用此 API。 - **浏览器自动化:**对于复杂情况,使用 Selenium 或 Playwright 等工具。这些工具控制真实的网络浏览器,可以执行 JavaScript 并在完全渲染页面后提取数据。这会消耗更多资源,但可以处理高度动态的网站。
网站也可能采用反爬虫技术。礼貌地行事,模仿浏览器行为(例如,User-Agent、头部),处理 CAPTCHA(通常需要手动干预或专门的服务),以及轮换代理有时会有帮助,但始终优先考虑道德行为。