Skip to content

Requests - 使用 Requests 进行网页抓取

Requests - 使用 Requests 和 BeautifulSoup 进行网页抓取

Section titled “Requests - 使用 Requests 和 BeautifulSoup 进行网页抓取”

网页抓取(Web scraping)是自动从网站提取数据的过程。requests 库是一个强大的工具,用于获取网页的原始 HTML 内容。然而,要解析此 HTML(即导航其结构并提取特定信息),您通常会使用专门的解析库。BeautifulSoup 是一个非常流行的 Python 库,用于执行此任务,它可以与 requests 很好地协同工作。

在开始任何网页抓取项目之前,理解并尊重道德和法律界限至关重要:

  • 检查 robots.txt:大多数网站都提供一个 robots.txt 文件(例如,domain.com/robots.txt),该文件指定了针对网络爬虫和机器人的规则。始终尊重这些规则。
  • 速率限制:不要在短时间内向服务器发送过多请求。这可能导致服务器过载,并可能导致您的 IP 被封锁。在请求之间实施延迟(例如,time.sleep())。
  • 标识您的爬虫:在请求中设置描述性的 User-Agent 头部。这有助于网站管理员识别您的抓取程序,并在必要时与您联系(例如,User-Agent: MyCoolScraper/1.0 (+http://mywebsite.com/scraper_info))。
  • 查阅服务条款:许多网站在其服务条款中明确说明了数据抓取政策。确保您的抓取活动符合规定。
  • 数据隐私和版权:注意个人数据和版权限制。仅抓取公开可用的数据并负责任地使用它们。
  • 轻量抓取:仅请求您需要的数据,并考虑是否有官方 API 可用,这始终是抓取的更好替代方案。

负责任的抓取是维持访问和良好网络公民身份的关键。

要开始抓取,您需要安装 requests 和 beautifulsoup4:

pip install requests beautifulsoup4

BeautifulSoup 可以使用不同的 HTML 解析器。Python 内置的 html.parser 是一个不错的默认选项。为了可能更快的解析速度和对格式错误的 HTML 更高的容错性,您可以安装 lxml:

pip install lxml

如果安装了 lxml,在创建 Soup 对象时指定 'lxml',BeautifulSoup 可以自动使用它。

典型的网页抓取过程包括以下步骤:

  1. 使用 requests.get(url, headers=...) 下载目标网页的 HTML 内容。
  2. 检查 response.status_code 以确保请求成功(例如,200 OK)。使用 response.raise_for_status() 进行简单的错误检查。
  3. 从 response.text 创建一个 BeautifulSoup 对象,指定解析器(例如,'html.parser' 或 'lxml')。
  4. 使用 BeautifulSoup 的方法(如 find()、find_all()、CSS 选择器)在解析后的 HTML 树中定位和提取所需的数据元素。

示例:抓取 ‘quotes.toscrape.com’ 的引言

Section titled “示例:抓取 ‘quotes.toscrape.com’ 的引言”

quotes.toscrape.com 是一个专为练习网页抓取而设计的网站。我们将编写一个脚本来提取引言、其作者和标签。

import requests
from bs4 import BeautifulSoup
import time # For respectful scraping
# 为了礼貌地进行抓取
URL = 'http://quotes.toscrape.com/'
# Define a user-agent to be polite
# 定义一个 User-Agent,以示友好
HEADERS = {
'User-Agent': 'MyTutorialScraper/1.0 (contact@example.com; +http://example.com/about_scraper)'
}
try:
print(f'Fetching URL: {URL}')
print(f'正在抓取 URL: {URL}')
response = requests.get(URL, headers=HEADERS, timeout=10) # Always use a timeout
# 始终使用超时设置
response.raise_for_status() # Raise an exception for bad status codes (4xx or 5xx)
# 对于不成功的状态码(4xx 或 5xx)引发异常
# Parse the HTML content using BeautifulSoup
# 使用 BeautifulSoup 解析 HTML 内容
# Using 'lxml' if available, otherwise 'html.parser'
# 如果 lxml 可用则使用它,否则使用 html.parser
try:
soup = BeautifulSoup(response.text, 'lxml')
except ImportError:
soup = BeautifulSoup(response.text, 'html.parser')
# Extract the page title (as an example)
# 提取页面标题(作为示例)
page_title = soup.title.string if soup.title else 'No title found'
page_title = soup.title.string if soup.title else '未找到标题'
print(f'Page Title: {page_title}\n')
print(f'页面标题: {page_title}\n')
# Find all quote containers (each quote is in a <div> with class='quote')
# 查找所有引言容器(每个引言都在一个 class='quote' 的 <div> 中)
quote_elements = soup.find_all('div', class_='quote')
if not quote_elements:
print('No quotes found. The website structure might have changed or content is not as expected.')
print('未找到引言。网站结构可能已更改或内容与预期不符。')
else:
print(f'Found {len(quote_elements)} quotes on the page. Displaying first few:')
print(f'页面上找到 {len(quote_elements)} 条引言。显示前几条:')
for i, quote_div in enumerate(quote_elements[:3]): # Display first 3 quotes for brevity
# 为简洁起见,显示前 3 条引言
text_element = quote_div.find('span', class_='text')
author_element = quote_div.find('small', class_='author')
tags_container = quote_div.find('div', class_='tags')
quote_text = text_element.get_text(strip=True) if text_element else 'N/A'
quote_author = author_element.get_text(strip=True) if author_element else 'N/A'
tags_list = []
if tags_container:
tag_elements = tags_container.find_all('a', class_='tag')
tags_list = [tag.get_text(strip=True) for tag in tag_elements]
print(f'\nQuote {i+1}:')
print(f'\n引言 {i+1}:')
print(f' Text: "{quote_text}"')
print(f' Author: {quote_author}')
print(f' Tags: {", ".join(tags_list) if tags_list else "N/A"}')
print(f' 文本: “{quote_text}”')
print(f' 作者: {quote_author}')
print(f' 标签: {", ".join(tags_list) if tags_list else "N/A"}')
# Respectful delay if scraping multiple pages/items
# time.sleep(1)
# 如果抓取多个页面/项目,请设置礼貌的延迟
except requests.exceptions.HTTPError as http_err:
print(f'HTTP error occurred while fetching {URL}: {http_err}')
print(f'抓取 {URL} 时发生 HTTP 错误: {http_err}')
except requests.exceptions.ConnectionError as conn_err:
print(f'Connection error occurred while fetching {URL}: {conn_err}')
print(f'抓取 {URL} 时发生连接错误: {conn_err}')
except requests.exceptions.Timeout as timeout_err:
print(f'Timeout error occurred while fetching {URL}: {timeout_err}')
print(f'抓取 {URL} 时发生超时错误: {timeout_err}')
except requests.exceptions.RequestException as req_err:
print(f'An general error occurred during the request to {URL}: {req_err}')
print(f'请求 {URL} 期间发生一般错误: {req_err}')
except Exception as e:
print(f'An unexpected error occurred during scraping: {e}')
print(f'抓取期间发生意外错误: {e}')

预期输出(根据当前网站内容和显示的引言数量可能略有不同):

Section titled “预期输出(根据当前网站内容和显示的引言数量可能略有不同):”
Fetching URL: http://quotes.toscrape.com/
Page Title: Quotes to Scrape
Found 10 quotes on the page. Displaying first few:
Quote 1:
Text: "The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking."
Author: Albert Einstein
Tags: change, deep-thoughts, thinking, world
Quote 2:
Text: "It is our choices, Harry, that show what we truly are, far more than our abilities."
Author: J.K. Rowling
Tags: abilities, choices
Quote 3:
Text: "There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle."
Author: Albert Einstein
Tags: inspirational, life, live, miracle, miracles
(Actual quotes and tags on the site may change.)
正在抓取 URL: http://quotes.toscrape.com/
页面标题: Quotes to Scrape
页面上找到 10 条引言。显示前几条:
引言 1:
文本: “The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”
作者: Albert Einstein
标签: change, deep-thoughts, thinking, world
引言 2:
文本: “It is our choices, Harry, that show what we truly are, far more than our abilities.”
作者: J.K. Rowling
标签: abilities, choices
引言 3:
文本: “There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”
作者: Albert Einstein
标签: inspirational, life, live, miracle, miracles
(网站上的实际引言和标签可能会发生变化。)

常用的 BeautifulSoup 数据提取操作

Section titled “常用的 BeautifulSoup 数据提取操作”
  • soup.find(‘tag_name’, attrs={‘attribute_name’: ‘value’}, class_=‘css_class_name’): 查找第一个匹配指定条件的元素。
  • soup.find_all(‘tag_name’, class_=‘css_class_name’, limit=N): 查找所有匹配的元素(如果指定了 limit,则最多查找 N 个),并将它们作为列表 (ResultSet) 返回。
  • element.get_text(strip=True, separator=’ ’): 从元素及其子元素中提取所有人类可读的文本。strip=True 删除每段文本的首尾空格。separator 可以用于连接文本片段。
  • element[‘attribute_name’](例如,link_tag['href']):访问元素的属性值。
  • CSS 选择器:soup.select(‘div.content > p’) 允许使用 CSS 选择器语法查找元素,这非常强大。例如,soup.select_one('#main-id .item-class') 查找匹配此 CSS 选择器的第一个元素。
  • 遍历树:.parent, .children(迭代器), .find_next_sibling(), .find_previous_sibling() 等,允许遍历解析后的 HTML 结构。

有关 BeautifulSoup 功能的全面指南,请参阅其官方文档。

许多现代网站在初始页面加载后使用 JavaScript 动态加载内容。由于 requests(以及由 requests 提供数据的 BeautifulSoup)只抓取初始 HTML 源代码,并且不执行 JavaScript,您可能会发现所需的数据缺失。

处理动态内容的策略包括:

  • **检查网络请求:**使用浏览器的开发者工具(Network 选项卡)查看 JavaScript 是否向 API 端点发起了 XHR/Fetch 请求,这些请求返回数据(通常是 JSON 格式)。如果是这样,您通常可以使用 requests 直接调用此 API。
  • **浏览器自动化:**对于复杂情况,使用 Selenium 或 Playwright 等工具。这些工具控制真实的网络浏览器,可以执行 JavaScript 并在完全渲染页面后提取数据。这会消耗更多资源,但可以处理高度动态的网站。

网站也可能采用反爬虫技术。礼貌地行事,模仿浏览器行为(例如,User-Agent、头部),处理 CAPTCHA(通常需要手动干预或专门的服务),以及轮换代理有时会有帮助,但始终优先考虑道德行为。