认识爬虫:如何使用 requests 模块模拟浏览器请求爬取网页信息?
'# 认识爬虫:如何使用 requests 模块模拟浏览器请求爬取网页信息?
一、背景与问题
在现代 Web 开发中,爬虫技术是数据采集的重要手段。无论是构建数据仓库、实现价格监控系统,还是进行市场分析,爬虫都扮演着关键角色。然而,传统浏览器请求和爬虫请求存在本质差异:浏览器会发送完整的 HTTP 头信息(如 User-Agent、Accept-Language 等),而简单的 requests 请求可能因缺少这些信息被服务器识别为非人类请求,从而触发反爬机制。
本文将深入解析 requests 模块的工作原理,结合真实开发场景,展示如何通过模拟浏览器行为安全地爬取网页信息。
二、基本原理
1. HTTP 协议基础
HTTP 是客户端与服务器通信的协议,其核心是请求-响应模型。当使用 requests 发起请求时,实际上是构建一个 HTTP 请求报文,包含以下要素:
- 请求方法(GET/POST/PUT/DELETE)
- 请求头(Headers):包含 User-Agent、Accept、Referer 等关键字段
- 请求体(Body):仅在 POST/PUT 等方法中存在
- 请求路径(Path):URL 的路径部分
服务器收到请求后,会根据规则返回响应报文,包含状态码(如 200 OK/403 Forbidden/503 Service Unavailable)和响应体(HTML 内容/JSON 数据等)。
2. requests 的工作原理
requests 是 Python 中最受欢迎的 HTTP 库,其底层依赖 urllib3,通过以下机制模拟浏览器行为:
- 自动处理重定向:自动跟随 301/302 状态码的跳转链接
- 会话管理:通过 Session 对象维护 cookies 和 headers
- 连接池:复用 TCP 连接提升性能
- 异常处理:内置连接超时、HTTP 错误码处理机制
三、环境准备
1. 安装依赖
pip install requests2. 开发环境
- Python 3.8+
- 建议使用虚拟环境(venv)隔离依赖
- 可选:配合 requests-cache 或 fake-useragent 等辅助库
四、核心实现
1. 基础 GET 请求
import requests
# 发起GET请求
response = requests.get('https://example.com')
# 打印响应状态码
print(f"Status Code: {response.status_code}")
# 打印响应内容
print(response.text)关键代码解析:
requests.get()构造了一个 HTTP GET 请求- 默认会发送
User-Agent: Python-requests/2.x.x的头信息 - 响应对象包含
status_code(HTTP 状态码)、text(响应内容)等属性
2. 模拟浏览器头信息
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
'Accept-Language': 'zh-CN,zh;q=0.9',
'Referer': 'https://www.google.com/'
}
response = requests.get('https://httpbin.org/headers', headers=headers)
print(response.json())关键代码解析:
- 自定义
User-Agent模拟 Chrome 浏览器 Referer字段用于告知服务器请求来源httpbin.org是测试用的 API 网站,返回请求头信息
3. 处理响应内容
if response.status_code == 200:
# 解析HTML内容
from bs4 import BeautifulSoup
soup = BeautifulSoup(response.text, 'html.parser')
print(soup.title.string)
else:
print(f"请求失败: {response.status_code}")关键代码解析:
- 使用
BeautifulSoup解析 HTML 文本 html.parser是 Python 内置的解析器- 需要安装
beautifulsoup4依赖(pip install beautifulsoup4)
五、完整案例
案例:爬取豆瓣图书Top250信息
import requests
from bs4 import BeautifulSoup
import time
def get_books(page):
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
'Referer': 'https://book.douban.com/'
}
url = f'https://book.douban.com/top250?start={page * 25}&filter= ' # 每页25条数据
response = requests.get(url, headers=headers)
if response.status_code != 200:
print(f"请求失败: {response.status_code}")
return []
soup = BeautifulSoup(response.text, 'html.parser')
items = soup.find_all('div', class_='item')
books = []
for item in items:
title = item.find('span', class_='title').text.strip()
author = item.find('div', class_='info').find('p').text.strip()
rating = item.find('span', class_='rating_num').text.strip()
books.append({
'title': title,
'author': author,
'rating': rating
})
return books
# 爬取前10页数据
all_books = []
for page in range(4): # 前4页共100条数据
print(f"正在爬取第 {page+1} 页...")
books = get_books(page)
all_books.extend(books)
time.sleep(1) # 模拟人类操作间隔
# 输出结果
for book in all_books[:10]:
print(f"书名: {book['title']}, 作者: {book['author']}, 评分: {book['rating']}")关键代码解析:
- 豆瓣图书Top250页面通过
start参数分页 - 每页包含25条数据,需爬取前4页获取100条数据
- 使用
time.sleep(1)模拟人类操作间隔,避免触发反爬机制 - 使用
BeautifulSoup提取书籍标题、作者、评分信息
六、源码解析
1. requests.get() 的实现原理
requests.get() 实际上调用了 requests.Session().get(),其核心流程如下:
def get(self, url, **kwargs):
return self.request('GET', url, **kwargs)- 构造 HTTP GET 请求报文
- 设置默认 headers(包含 User-Agent 等)
- 发送请求并处理响应
- 自动处理重定向(可配置
allow_redirects参数)
2. 会话管理(Session)
session = requests.Session()
session.headers.update({
'User-Agent': 'Custom User Agent',
'Accept-Encoding': 'gzip, deflate'
})
response = session.get('https://example.com')- Session 对象可以持久化 cookies
- 可以设置全局 headers,避免重复配置
- 支持添加代理、验证证书等高级功能
七、进阶使用
1. 处理 cookies
cookies = {
'session_id': '123456',
'user_token': 'abcdefg'
}
response = requests.get('https://example.com', cookies=cookies)2. 设置超时和重试
response = requests.get(
'https://example.com',
timeout=5, # 设置超时时间
allow_redirects=False # 禁用重定向
)3. 使用代理服务器
proxies = {
'http': 'http://10.10.1.10:3128',
'https': 'http://10.10.1.10:1080'
}
response = requests.get('https://example.com', proxies=proxies)八、性能与工程实践
1. 并发请求优化
from concurrent.futures import ThreadPoolExecutor
def fetch_page(page):
# 实现爬取逻辑
return f"Page {page} data"
with ThreadPoolExecutor(max_workers=5) as executor:
results = list(executor.map(fetch_page, range(10)))2. 使用缓存减少请求
import requests_cache
requests_cache.install_cache('douban_cache', expire_after=3600) # 缓存1小时
response = requests.get('https://example.com')3. 异常处理机制
try:
response = requests.get('https://example.com', timeout=5)
response.raise_for_status() # 如果响应状态码不是200,抛出异常
except requests.exceptions.RequestException as e:
print(f"请求异常: {e}")九、常见问题与踩坑
1. 常见错误
| 错误类型 | 原因 | 解决方案 |
|---|---|---|
| 403 Forbidden | 未正确设置 headers | 添加 User-Agent、Referer 等字段 |
| 503 Service Unavailable | 服务器暂时不可用 | 增加重试机制,设置 timeout |
| 429 Too Many Requests | 被限速 | 增加请求间隔,使用代理服务器 |
| 10054 连接被拒绝 | 服务器主动断开 | 检查防火墙设置,更换代理 |
2. 常见问题分析
问题:爬虫被封IP
- 原因:短时间内发送大量请求,触发服务器限流机制
- 解决方案:增加请求间隔,使用代理池,设置
headers模拟真实用户
问题:无法解析响应内容
- 原因:服务器返回的是二进制数据(如图片)而非 HTML
- 解决方案:检查响应
content-type,使用response.content获取原始数据
问题:请求超时
- 原因:网络不稳定或服务器处理时间过长
- 解决方案:设置合理的
timeout值,增加重试机制
十、最佳实践
1. 推荐实践方案
| 场景 | 推荐方案 | 原因 |
|---|---|---|
| 简单数据采集 | requests + BeautifulSoup | 简单易用,适合静态页面 |
| 复杂交互 | Selenium | 可模拟真实浏览器行为 |
| 高并发爬取 | asyncio + aiohttp | 非阻塞IO,提升性能 |
| 需要验证 | requests + PyQuery | 结合 CSS 选择器提高解析效率 |
2. 推荐配置参数
headers:设置完整的 User-Agent、Referer 等字段timeout:设置合理超时时间(建议 3-5 秒)proxies:使用代理服务器避免IP被封verify:验证SSL证书(生产环境建议启用)
十一、总结
requests 模块是 Python 中最常用的 HTTP 请求库,其通过模拟浏览器行为实现网页数据采集。本文深入解析了其工作原理,结合真实开发场景展示了如何通过合理设置 headers、处理响应内容、优化性能等手段实现高效爬虫。
在实际项目中,requests 适用于静态页面数据采集、简单的 API 调用等场景。但需注意:对于需要复杂交互的网页(如 JavaScript 渲染内容)、有严格反爬机制的网站(如电商平台),应考虑使用 Selenium、Playwright 等工具。同时,需遵守目标网站的 robots.txt 规则,避免对服务器造成过大负担。
通过合理配置 headers、添加异常处理、优化请求频率等措施,可以显著提高爬虫的稳定性和安全性。在开发过程中,始终要保持对技术原理的理解,才能应对各种实际问题和挑战。
评论已关闭