Skip to content

处理 HTTP 请求响应

当你使用 requests 库发起请求时,服务器会返回一个 HTTP 响应。requests 返回的 Response 对象包含了丰富的信息,例如响应体、状态码、头部(headers)等。本章将探讨如何访问和处理这些数据。

  • 访问响应内容(文本、字节、JSON)
  • 检查状态码
  • 检查头部(Headers)
  • 处理编码
  • 处理原始(Raw)和二进制(Binary)响应

让我们向一个公共 API 发起请求,演示如何处理响应:

import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')

response.text 属性提供了响应体作为 Unicode 字符串。requests 会根据 HTTP 头部(例如带有 charset 参数的 Content-Type)自动解码内容。

import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')
print("Response Text:")
print(response.text)
Response Text:
{
"userId": 1,
"id": 1,
"title": "delectus aut autem",
"completed": false
}

这类似于在浏览器中查看网页的“源代码”(如果内容是 HTML),或者原始的 JSON 字符串(如果它是一个 JSON API)。

response.content 属性提供了响应体作为原始字节。这对于非文本数据(如图像、PDF)或你需要自己处理解码的情况非常有用。

import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')
print("\nResponse Content (Bytes):")
print(response.content)
Response Content (Bytes):
b'{\n "userId": 1,\n "id": 1,\n "title": "delectus aut autem",\n "completed": false\n}'

注意 b'' 前缀,它表示这是一个字节字符串。这是基于字符集进行任何解码之前的原始数据。

如果响应内容是 JSON 格式,你可以使用 response.json() 方法将其直接解析成 Python 字典或列表。

import requests
import json # 用于处理可能的 JSONDecodeError
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')
try:
data = response.json()
print("\nJSON Data:")
print(data)
print(f"Type of parsed data: {type(data)}")
print(f"Title from JSON: {data['title']}")
except json.JSONDecodeError:
print("\nFailed to decode JSON.")
JSON Data:
{'userId': 1, 'id': 1, 'title': 'delectus aut autem', 'completed': False}
Type of parsed data: <class 'dict'>
Title from JSON: delectus aut autem

如果响应不是有效的 JSON,response.json() 将会抛出 json.JSONDecodeError 异常(或者在较新版本的 requests 中是 requests.exceptions.JSONDecodeError,它是前者的子类)。

response.status_code 属性提供了响应的 HTTP 状态码。检查请求是否成功(例如 200 OK)或遇到错误(例如 404 Not Found, 500 Internal Server Error)时,它至关重要。

import requests
response_ok = requests.get('https://jsonplaceholder.typicode.com/todos/1')
print(f"Status for /todos/1: {response_ok.status_code}")
response_not_found = requests.get('https://jsonplaceholder.typicode.com/nonexistent')
print(f"Status for /nonexistent: {response_not_found.status_code}")
# 最佳实践:对不良状态码(4xx 或 5xx)抛出异常
try:
response_not_found.raise_for_status()
except requests.exceptions.HTTPError as e:
print(f"HTTP error occurred for /nonexistent: {e}")
Status for /todos/1: 200
Status for /nonexistent: 404
HTTP error occurred for /nonexistent: 404 Client Error: Not Found for url: https://jsonplaceholder.typicode.com/nonexistent

response.raise_for_status() 是一种方便的方式,用于断言请求是否成功。

HTTP 头部(headers)提供有关响应的元数据。你可以通过 response.headers 访问它们。它是一个不区分大小写的类似字典的对象。

import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')
print("\nResponse Headers:")
for key, value in response.headers.items():
print(f" {key}: {value}")
print(f"\nContent-Type Header: {response.headers['Content-Type']}")
# Access is case-insensitive
print(f"content-type Header (lowercase key): {response.headers['content-type']}")
Response Headers:
Date: ...
Content-Type: application/json; charset=utf-8
Content-Length: ...
Connection: keep-alive
...
Content-Type Header: application/json; charset=utf-8
content-type Header (lowercase key): application/json; charset=utf-8

requests 尝试根据头部猜测响应内容的编码。你可以通过 response.encoding 来检查它。

import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')
print(f"\nGuessed Encoding: {response.encoding}")
# 如果 requests 猜测不正确,或者没有指定字符集,
# 你可以在访问 .text 之前手动设置编码
# response.encoding = 'ISO-8859-1' # 示例
# print(f"手动设置的编码: {response.encoding}")
# print(response.text) # .text 现在将使用新的编码
Guessed Encoding: utf-8

原始(Raw)和二进制(Binary)响应

Section titled “原始(Raw)和二进制(Binary)响应”

要对响应流进行底层访问,你可以使用 response.raw。这要求在你的请求中设置 stream=True。response.raw 是一个类似文件的对象(具体来说,是一个 urllib3.response.HTTPResponse 实例)。

import requests
# 示例:分块下载一个小型图像
# 使用一个图像 URL 的占位符,如果需要请替换为真实的 URL
# 为了演示,我们将使用一个支持流式传输的文本接口
response = requests.get('https://jsonplaceholder.typicode.com/todos/1', stream=True)
print(f"\nRaw response object: {response.raw}")
# 你可以从原始流中读取数据
# 使用完 raw 后释放连接是一个好习惯
with response:
try:
# 读取原始响应的一小块数据
raw_chunk = response.raw.read(50) # 读取最多 50 字节
print(f"First 50 bytes of raw response: {raw_chunk}")
# 在实际场景中,你可能会将数据分块写入文件
# for chunk in response.iter_content(chunk_size=8192):
# if chunk: # 过滤掉 keep-alive 新块
# # 处理块 (例如 f.write(chunk))
# pass
except Exception as e:
print(f"Error reading raw stream: {e}")

预期输出(原始响应 - 字节内容将与 JSON 内容匹配)

Section titled “预期输出(原始响应 - 字节内容将与 JSON 内容匹配)”
Raw response object: <urllib3.response.HTTPResponse object at 0x...>
First 50 bytes of raw response: b'{\n "userId": 1,\n "id": 1,\n "title": "delectus aut au'

流式传输(stream=True)对于下载大文件或处理流式 API 特别有用,在这种情况下,你不想一次性将整个响应加载到内存中。

二进制响应(回顾:response.content)

Section titled “二进制响应(回顾:response.content)”

如前所述,response.content 以字节形式提供响应体。这是获取二进制数据(如图像、音频文件或任何非文本内容)最直接的方式。当整个内容可以轻松载入内存时,它非常适用。

import requests
# 假设的图像 URL
# response = requests.get('https://via.placeholder.com/150.png')
# with open('image.png', 'wb') as f:
# f.write(response.content)
# print("\n图像已下载为 image.png (使用 response.content)")
# 对于本示例,让我们再次展示 .content 与 JSON API 的用法:
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')
print("\nBinary content (response.content):")
print(response.content)
Binary content (response.content):
b'{\n "userId": 1,\n "id": 1,\n "title": "delectus aut autem",\n "completed": false\n}'

理解如何处理 HTTP 响应的不同方面是构建与 Web 服务交互的健壮应用程序的关键。更多详情,请参阅官方文档中的Requests 快速入门和高级用法部分。