Python进阶--查询商品历史价格(基于慢慢买比价网的爬虫)
'# Python进阶--查询商品历史价格(基于慢慢买比价网的爬虫)
一、背景与问题
在电商价格监控、市场分析、竞品研究等场景中,获取商品历史价格数据是核心需求。以"慢慢买比价网"为例,其网页结构包含商品历史价格走势图、价格波动趋势等关键信息,但这些数据通常需要通过爬虫技术获取。
当前面临的核心问题包括:
- 网站反爬机制的应对
- 多页数据的高效抓取
- 历史价格数据的结构化存储
- 大数据量下的性能优化
- 合法合规的爬虫策略
二、基本原理
1. 网站结构分析
通过浏览器开发者工具分析慢慢买比价网的页面结构,发现:
- 商品价格数据存储在
<div class="price-chart">容器中 - 每个价格点包含
<div class="price-item">,包含date和price两个子节点 - 分页信息位于
<div class="pagination">中,包含<a>标签的data-page属性
2. 爬虫流程
- 发送HTTP请求获取网页内容
- 解析HTML提取价格数据
- 处理分页获取多页数据
- 存储数据到数据库或文件
- 异常处理与重试机制
3. 技术挑战
- 动态加载内容(需处理AJAX请求)
- 防止IP被封禁(需设置请求头、使用代理)
- 数据清洗与去重(处理异常数据格式)
三、环境准备
1. 必备工具
- Python 3.9+
- requests: HTTP请求库
- BeautifulSoup: HTML解析库
- pandas: 数据处理库
- sqlite3: 轻量级数据库
- fake-useragent: 随机User-Agent库
pip install requests beautifulsoup4 pandas fake-useragent四、核心实现
1. 请求头构造
from fake_useragent import UserAgent
def get_headers():
ua = UserAgent()
return {
"User-Agent": ua.random,
"Accept-Language": "en-US,en;q=0.9",
"Referer": "https://www.maimaibiao.com"
}关键点:
- 随机User-Agent防止被识别为爬虫
- 设置Referer避免被服务器拒绝
- 需处理网站反爬机制(如IP限速)
2. 分页数据抓取
import requests
from bs4 import BeautifulSoup
def fetch_page(page_number):
url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
headers = get_headers()
response = requests.get(url, headers=headers, timeout=10)
return BeautifulSoup(response.text, 'html.parser')注意:
- 需替换
product_id为实际商品ID - 需处理可能的403/429响应
- 建议添加重试机制
3. 数据解析与存储
def parse_page(soup):
price_data = []
for item in soup.select('.price-item'):
date = item.select_one('.date').text.strip()
price = item.select_one('.price').text.strip()
price_data.append({
'date': date,
'price': float(price.replace('¥', '').replace(',', '')),
'product_id': product_id
})
return price_data
def save_to_sqlite(data):
import sqlite3
conn = sqlite3.connect('price_history.db')
c = conn.cursor()
c.execute("CREATE TABLE IF NOT EXISTS prices (id INTEGER PRIMARY KEY, date TEXT, price REAL, product_id INTEGER)")
c.executemany("INSERT INTO prices (date, price, product_id) VALUES (?, ?, ?)", data)
conn.commit()
conn.close()五、完整案例
1. 爬虫主流程
def main():
product_id = "12345" # 替换为实际商品ID
max_pages = 10
all_data = []
for page in range(1, max_pages+1):
print(f"正在抓取第 {page} 页...")
soup = fetch_page(page)
if not soup.select('.price-item'):
break # 没有更多数据
data = parse_page(soup)
all_data.extend(data)
save_to_sqlite(all_data)
print(f"共抓取 {len(all_data)} 条价格数据")2. 优化措施
- 添加请求间隔(防止触发反爬)
- 使用代理IP池
- 增加异常处理
- 使用缓存机制
import time
import random
def fetch_page(page_number):
# ... 之前的代码 ...
time.sleep(random.uniform(0.5, 1.5)) # 随机间隔六、源码解析
1. 网络请求模块
def get_headers():
# 随机User-Agent构造
ua = UserAgent()
return {
"User-Agent": ua.random,
"Accept-Language": "en-US,en;q=0.9",
"Referer": "https://www.maimaibiao.com"
}解析:
- 使用fake-useragent库生成随机User-Agent
- 设置合理的HTTP头字段
- Referer字段模拟正常浏览器访问
2. 数据解析模块
def parse_page(soup):
price_data = []
for item in soup.select('.price-item'):
date = item.select_one('.date').text.strip()
price = item.select_one('.price').text.strip()
price_data.append({
'date': date,
'price': float(price.replace('¥', '').replace(',', '')),
'product_id': product_id
})
return price_data关键点:
- 使用CSS选择器高效提取数据
- 处理价格字符串格式
- 结构化存储数据
七、进阶使用
1. 异步爬虫优化
使用aiohttp实现异步请求:
import aiohttp
import asyncio
async def fetch_page(session, page_number):
url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
headers = get_headers()
async with session.get(url, headers=headers) as response:
return await response.text()优势:
- 提高爬取效率(可同时处理多个请求)
- 更适合大规模数据抓取
- 需配合async/await使用
2. 数据持久化优化
使用SQLite的批量插入:
def save_to_sqlite(data):
import sqlite3
conn = sqlite3.connect('price_history.db')
c = conn.cursor()
c.execute("CREATE TABLE IF NOT EXISTS prices (id INTEGER PRIMARY KEY, date TEXT, price REAL, product_id INTEGER)")
c.executemany("INSERT INTO prices (date, price, product_id) VALUES (?, ?, ?)", data)
conn.commit()
conn.close()八、性能与工程实践
1. 性能优化策略
- 使用缓存(本地缓存+网络缓存)
- 增加并发控制(限制请求频率)
- 使用数据库索引优化查询
- 压缩传输数据(使用gzip)
- 避免不必要的数据处理
2. 异常处理机制
def fetch_page(page_number):
try:
url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
headers = get_headers()
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
return BeautifulSoup(response.text, 'html.parser')
except requests.exceptions.RequestException as e:
print(f"请求失败: {e}")
return None3. 安全风险分析
- 频繁请求可能导致IP被封禁
- 网站可能增加验证码验证
- 数据可能包含敏感信息
- 需遵守robots.txt协议
九、常见问题与踩坑
1. 常见错误
- Error 403 Forbidden:未正确设置headers
- TimeoutError:网络不稳定或服务器响应慢
- ParseError:网页结构变更
- Duplicate Data:未处理重复数据
2. 错误解决办法
- 添加
User-Agent和Referer头 - 增加重试机制
- 使用正则表达式或XPath处理动态内容
- 添加数据去重逻辑
3. 典型问题示例
# 错误示例:未处理异常
def fetch_page(page_number):
url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
response = requests.get(url)
return BeautifulSoup(response.text, 'html.parser')改进:
# 改进示例:添加异常处理
def fetch_page(page_number):
try:
url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
headers = get_headers()
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
return BeautifulSoup(response.text, 'html.parser')
except requests.exceptions.RequestException as e:
print(f"请求失败: {e}")
return None十、最佳实践
1. 推荐方案
- 使用异步爬虫提高效率
- 配合代理IP池防止被封
- 使用数据库索引优化查询
- 添加日志记录和监控
- 定期更新爬虫逻辑(应对网站结构变更)
2. 使用场景
- 价格监控系统
- 市场趋势分析
- 竞品对比研究
- 数据可视化展示
3. 不推荐使用场景
- 未取得网站授权
- 网站明确禁止爬虫
- 需要处理大量动态内容(需模拟浏览器)
- 存在严重法律风险
十一、总结
本文深入探讨了基于慢慢买比价网的商品历史价格爬虫技术,从原理分析、代码实现到性能优化进行了全面解析。重点包括:
- 如何应对网站反爬机制
- 分页数据的高效抓取方法
- 数据结构化存储方案
- 性能优化策略
- 安全风险分析
实际开发中,建议根据具体需求选择合适的方案,注意遵守相关法律法规。对于需要长期运行的爬虫系统,建议采用分布式架构和监控机制,确保数据的准确性和系统的稳定性。
评论已关闭