处理 HTTP 请求响应
Requests - 处理 HTTP 响应数据
Section titled “Requests - 处理 HTTP 响应数据”当你使用 requests 库发起请求时,服务器会返回一个 HTTP 响应。requests 返回的 Response 对象包含了丰富的信息,例如响应体、状态码、头部(headers)等。本章将探讨如何访问和处理这些数据。
- 访问响应内容(文本、字节、JSON)
- 检查状态码
- 检查头部(Headers)
- 处理编码
- 处理原始(Raw)和二进制(Binary)响应
访问响应内容
Section titled “访问响应内容”让我们向一个公共 API 发起请求,演示如何处理响应:
import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')文本内容 (response.text)
Section titled “文本内容 (response.text)”response.text 属性提供了响应体作为 Unicode 字符串。requests 会根据 HTTP 头部(例如带有 charset 参数的 Content-Type)自动解码内容。
import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')print("Response Text:")print(response.text)预期输出(文本)
Section titled “预期输出(文本)”Response Text:{ "userId": 1, "id": 1, "title": "delectus aut autem", "completed": false}这类似于在浏览器中查看网页的“源代码”(如果内容是 HTML),或者原始的 JSON 字符串(如果它是一个 JSON API)。
字节内容 (response.content)
Section titled “字节内容 (response.content)”response.content 属性提供了响应体作为原始字节。这对于非文本数据(如图像、PDF)或你需要自己处理解码的情况非常有用。
import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')print("\nResponse Content (Bytes):")print(response.content)预期输出(字节)
Section titled “预期输出(字节)”Response Content (Bytes):b'{\n "userId": 1,\n "id": 1,\n "title": "delectus aut autem",\n "completed": false\n}'注意 b'' 前缀,它表示这是一个字节字符串。这是基于字符集进行任何解码之前的原始数据。
JSON 响应 (response.json())
Section titled “JSON 响应 (response.json())”如果响应内容是 JSON 格式,你可以使用 response.json() 方法将其直接解析成 Python 字典或列表。
import requestsimport json # 用于处理可能的 JSONDecodeError
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')try: data = response.json() print("\nJSON Data:") print(data) print(f"Type of parsed data: {type(data)}") print(f"Title from JSON: {data['title']}")except json.JSONDecodeError: print("\nFailed to decode JSON.")预期输出(JSON)
Section titled “预期输出(JSON)”JSON Data:{'userId': 1, 'id': 1, 'title': 'delectus aut autem', 'completed': False}Type of parsed data: <class 'dict'>Title from JSON: delectus aut autem如果响应不是有效的 JSON,response.json() 将会抛出 json.JSONDecodeError 异常(或者在较新版本的 requests 中是 requests.exceptions.JSONDecodeError,它是前者的子类)。
response.status_code 属性提供了响应的 HTTP 状态码。检查请求是否成功(例如 200 OK)或遇到错误(例如 404 Not Found, 500 Internal Server Error)时,它至关重要。
import requests
response_ok = requests.get('https://jsonplaceholder.typicode.com/todos/1')print(f"Status for /todos/1: {response_ok.status_code}")
response_not_found = requests.get('https://jsonplaceholder.typicode.com/nonexistent')print(f"Status for /nonexistent: {response_not_found.status_code}")
# 最佳实践:对不良状态码(4xx 或 5xx)抛出异常try: response_not_found.raise_for_status()except requests.exceptions.HTTPError as e: print(f"HTTP error occurred for /nonexistent: {e}")预期输出(状态码)
Section titled “预期输出(状态码)”Status for /todos/1: 200Status for /nonexistent: 404HTTP error occurred for /nonexistent: 404 Client Error: Not Found for url: https://jsonplaceholder.typicode.com/nonexistentresponse.raise_for_status() 是一种方便的方式,用于断言请求是否成功。
检查头部(Headers)
Section titled “检查头部(Headers)”HTTP 头部(headers)提供有关响应的元数据。你可以通过 response.headers 访问它们。它是一个不区分大小写的类似字典的对象。
import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')
print("\nResponse Headers:")for key, value in response.headers.items(): print(f" {key}: {value}")
print(f"\nContent-Type Header: {response.headers['Content-Type']}")# Access is case-insensitiveprint(f"content-type Header (lowercase key): {response.headers['content-type']}")预期输出(头部 - 已截断)
Section titled “预期输出(头部 - 已截断)”Response Headers: Date: ... Content-Type: application/json; charset=utf-8 Content-Length: ... Connection: keep-alive ...
Content-Type Header: application/json; charset=utf-8content-type Header (lowercase key): application/json; charset=utf-8requests 尝试根据头部猜测响应内容的编码。你可以通过 response.encoding 来检查它。
import requests
response = requests.get('https://jsonplaceholder.typicode.com/todos/1')print(f"\nGuessed Encoding: {response.encoding}")
# 如果 requests 猜测不正确,或者没有指定字符集,# 你可以在访问 .text 之前手动设置编码# response.encoding = 'ISO-8859-1' # 示例# print(f"手动设置的编码: {response.encoding}")# print(response.text) # .text 现在将使用新的编码预期输出(编码)
Section titled “预期输出(编码)”Guessed Encoding: utf-8原始(Raw)和二进制(Binary)响应
Section titled “原始(Raw)和二进制(Binary)响应”原始响应 (response.raw)
Section titled “原始响应 (response.raw)”要对响应流进行底层访问,你可以使用 response.raw。这要求在你的请求中设置 stream=True。response.raw 是一个类似文件的对象(具体来说,是一个 urllib3.response.HTTPResponse 实例)。
import requests
# 示例:分块下载一个小型图像# 使用一个图像 URL 的占位符,如果需要请替换为真实的 URL# 为了演示,我们将使用一个支持流式传输的文本接口response = requests.get('https://jsonplaceholder.typicode.com/todos/1', stream=True)
print(f"\nRaw response object: {response.raw}")
# 你可以从原始流中读取数据# 使用完 raw 后释放连接是一个好习惯with response: try: # 读取原始响应的一小块数据 raw_chunk = response.raw.read(50) # 读取最多 50 字节 print(f"First 50 bytes of raw response: {raw_chunk}") # 在实际场景中,你可能会将数据分块写入文件 # for chunk in response.iter_content(chunk_size=8192): # if chunk: # 过滤掉 keep-alive 新块 # # 处理块 (例如 f.write(chunk)) # pass except Exception as e: print(f"Error reading raw stream: {e}")预期输出(原始响应 - 字节内容将与 JSON 内容匹配)
Section titled “预期输出(原始响应 - 字节内容将与 JSON 内容匹配)”Raw response object: <urllib3.response.HTTPResponse object at 0x...>First 50 bytes of raw response: b'{\n "userId": 1,\n "id": 1,\n "title": "delectus aut au'流式传输(stream=True)对于下载大文件或处理流式 API 特别有用,在这种情况下,你不想一次性将整个响应加载到内存中。
二进制响应(回顾:response.content)
Section titled “二进制响应(回顾:response.content)”如前所述,response.content 以字节形式提供响应体。这是获取二进制数据(如图像、音频文件或任何非文本内容)最直接的方式。当整个内容可以轻松载入内存时,它非常适用。
import requests
# 假设的图像 URL# response = requests.get('https://via.placeholder.com/150.png')# with open('image.png', 'wb') as f:# f.write(response.content)# print("\n图像已下载为 image.png (使用 response.content)")
# 对于本示例,让我们再次展示 .content 与 JSON API 的用法:response = requests.get('https://jsonplaceholder.typicode.com/todos/1')print("\nBinary content (response.content):")print(response.content)预期输出(二进制内容)
Section titled “预期输出(二进制内容)”Binary content (response.content):b'{\n "userId": 1,\n "id": 1,\n "title": "delectus aut autem",\n "completed": false\n}'理解如何处理 HTTP 响应的不同方面是构建与 Web 服务交互的健壮应用程序的关键。更多详情,请参阅官方文档中的Requests 快速入门和高级用法部分。