Python进阶--查询商品历史价格(基于慢慢买比价网的爬虫)

'# Python进阶--查询商品历史价格(基于慢慢买比价网的爬虫)

一、背景与问题

在电商价格监控、市场分析、竞品研究等场景中,获取商品历史价格数据是核心需求。以"慢慢买比价网"为例,其网页结构包含商品历史价格走势图、价格波动趋势等关键信息,但这些数据通常需要通过爬虫技术获取。

当前面临的核心问题包括:

  1. 网站反爬机制的应对
  2. 多页数据的高效抓取
  3. 历史价格数据的结构化存储
  4. 大数据量下的性能优化
  5. 合法合规的爬虫策略

二、基本原理

1. 网站结构分析

通过浏览器开发者工具分析慢慢买比价网的页面结构,发现:

  • 商品价格数据存储在<div class="price-chart">容器中
  • 每个价格点包含<div class="price-item">,包含date和price两个子节点
  • 分页信息位于<div class="pagination">中,包含<a>标签的data-page属性

2. 爬虫流程

  1. 发送HTTP请求获取网页内容
  2. 解析HTML提取价格数据
  3. 处理分页获取多页数据
  4. 存储数据到数据库或文件
  5. 异常处理与重试机制

3. 技术挑战

  • 动态加载内容(需处理AJAX请求)
  • 防止IP被封禁(需设置请求头、使用代理)
  • 数据清洗与去重(处理异常数据格式)

三、环境准备

1. 必备工具

  • Python 3.9+
  • requests: HTTP请求库
  • BeautifulSoup: HTML解析库
  • pandas: 数据处理库
  • sqlite3: 轻量级数据库
  • fake-useragent: 随机User-Agent库
pip install requests beautifulsoup4 pandas fake-useragent

四、核心实现

1. 请求头构造

from fake_useragent import UserAgent

def get_headers():
    ua = UserAgent()
    return {
        "User-Agent": ua.random,
        "Accept-Language": "en-US,en;q=0.9",
        "Referer": "https://www.maimaibiao.com"
    }

关键点:

  • 随机User-Agent防止被识别为爬虫
  • 设置Referer避免被服务器拒绝
  • 需处理网站反爬机制(如IP限速)

2. 分页数据抓取

import requests
from bs4 import BeautifulSoup

def fetch_page(page_number):
    url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
    headers = get_headers()
    response = requests.get(url, headers=headers, timeout=10)
    return BeautifulSoup(response.text, 'html.parser')

注意:

  • 需替换product_id为实际商品ID
  • 需处理可能的403/429响应
  • 建议添加重试机制

3. 数据解析与存储

def parse_page(soup):
    price_data = []
    for item in soup.select('.price-item'):
        date = item.select_one('.date').text.strip()
        price = item.select_one('.price').text.strip()
        price_data.append({
            'date': date,
            'price': float(price.replace('¥', '').replace(',', '')),
            'product_id': product_id
        })
    return price_data

def save_to_sqlite(data):
    import sqlite3
    conn = sqlite3.connect('price_history.db')
    c = conn.cursor()
    c.execute("CREATE TABLE IF NOT EXISTS prices (id INTEGER PRIMARY KEY, date TEXT, price REAL, product_id INTEGER)")
    c.executemany("INSERT INTO prices (date, price, product_id) VALUES (?, ?, ?)", data)
    conn.commit()
    conn.close()

五、完整案例

1. 爬虫主流程

def main():
    product_id = "12345"  # 替换为实际商品ID
    max_pages = 10
    all_data = []
    
    for page in range(1, max_pages+1):
        print(f"正在抓取第 {page} 页...")
        soup = fetch_page(page)
        if not soup.select('.price-item'):
            break  # 没有更多数据
        
        data = parse_page(soup)
        all_data.extend(data)
    
    save_to_sqlite(all_data)
    print(f"共抓取 {len(all_data)} 条价格数据")

2. 优化措施

  • 添加请求间隔(防止触发反爬)
  • 使用代理IP池
  • 增加异常处理
  • 使用缓存机制
import time
import random

def fetch_page(page_number):
    # ... 之前的代码 ...
    time.sleep(random.uniform(0.5, 1.5))  # 随机间隔

六、源码解析

1. 网络请求模块

def get_headers():
    # 随机User-Agent构造
    ua = UserAgent()
    return {
        "User-Agent": ua.random,
        "Accept-Language": "en-US,en;q=0.9",
        "Referer": "https://www.maimaibiao.com"
    }

解析:

  • 使用fake-useragent库生成随机User-Agent
  • 设置合理的HTTP头字段
  • Referer字段模拟正常浏览器访问

2. 数据解析模块

def parse_page(soup):
    price_data = []
    for item in soup.select('.price-item'):
        date = item.select_one('.date').text.strip()
        price = item.select_one('.price').text.strip()
        price_data.append({
            'date': date,
            'price': float(price.replace('¥', '').replace(',', '')),
            'product_id': product_id
        })
    return price_data

关键点:

  • 使用CSS选择器高效提取数据
  • 处理价格字符串格式
  • 结构化存储数据

七、进阶使用

1. 异步爬虫优化

使用aiohttp实现异步请求:

import aiohttp
import asyncio

async def fetch_page(session, page_number):
    url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
    headers = get_headers()
    async with session.get(url, headers=headers) as response:
        return await response.text()

优势:

  • 提高爬取效率(可同时处理多个请求)
  • 更适合大规模数据抓取
  • 需配合async/await使用

2. 数据持久化优化

使用SQLite的批量插入:

def save_to_sqlite(data):
    import sqlite3
    conn = sqlite3.connect('price_history.db')
    c = conn.cursor()
    c.execute("CREATE TABLE IF NOT EXISTS prices (id INTEGER PRIMARY KEY, date TEXT, price REAL, product_id INTEGER)")
    c.executemany("INSERT INTO prices (date, price, product_id) VALUES (?, ?, ?)", data)
    conn.commit()
    conn.close()

八、性能与工程实践

1. 性能优化策略

  • 使用缓存(本地缓存+网络缓存)
  • 增加并发控制(限制请求频率)
  • 使用数据库索引优化查询
  • 压缩传输数据(使用gzip)
  • 避免不必要的数据处理

2. 异常处理机制

def fetch_page(page_number):
    try:
        url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
        headers = get_headers()
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        return BeautifulSoup(response.text, 'html.parser')
    except requests.exceptions.RequestException as e:
        print(f"请求失败: {e}")
        return None

3. 安全风险分析

  • 频繁请求可能导致IP被封禁
  • 网站可能增加验证码验证
  • 数据可能包含敏感信息
  • 需遵守robots.txt协议

九、常见问题与踩坑

1. 常见错误

  • Error 403 Forbidden:未正确设置headers
  • TimeoutError:网络不稳定或服务器响应慢
  • ParseError:网页结构变更
  • Duplicate Data:未处理重复数据

2. 错误解决办法

  • 添加User-Agent和Referer头
  • 增加重试机制
  • 使用正则表达式或XPath处理动态内容
  • 添加数据去重逻辑

3. 典型问题示例

# 错误示例:未处理异常
def fetch_page(page_number):
    url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
    response = requests.get(url)
    return BeautifulSoup(response.text, 'html.parser')

改进:

# 改进示例:添加异常处理
def fetch_page(page_number):
    try:
        url = f"https://www.maimaibiao.com/product/{product_id}/history?page={page_number}"
        headers = get_headers()
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        return BeautifulSoup(response.text, 'html.parser')
    except requests.exceptions.RequestException as e:
        print(f"请求失败: {e}")
        return None

十、最佳实践

1. 推荐方案

  • 使用异步爬虫提高效率
  • 配合代理IP池防止被封
  • 使用数据库索引优化查询
  • 添加日志记录和监控
  • 定期更新爬虫逻辑(应对网站结构变更)

2. 使用场景

  • 价格监控系统
  • 市场趋势分析
  • 竞品对比研究
  • 数据可视化展示

3. 不推荐使用场景

  • 未取得网站授权
  • 网站明确禁止爬虫
  • 需要处理大量动态内容(需模拟浏览器)
  • 存在严重法律风险

十一、总结

本文深入探讨了基于慢慢买比价网的商品历史价格爬虫技术,从原理分析、代码实现到性能优化进行了全面解析。重点包括:

  • 如何应对网站反爬机制
  • 分页数据的高效抓取方法
  • 数据结构化存储方案
  • 性能优化策略
  • 安全风险分析

实际开发中,建议根据具体需求选择合适的方案,注意遵守相关法律法规。对于需要长期运行的爬虫系统,建议采用分布式架构和监控机制,确保数据的准确性和系统的稳定性。

最后修改于:2026年09月22日 06:39

评论已关闭

推荐阅读

AIGC实战——Transformer模型
2024年12月01日
Socket TCP 和 UDP 编程基础(Python)
2024年11月30日
python , tcp , udp
如何使用 ChatGPT 进行学术润色?你需要这些指令
2024年12月01日
AI
最新 Python 调用 OpenAi 详细教程实现问答、图像合成、图像理解、语音合成、语音识别(详细教程)
2024年11月24日
ChatGPT 和 DALL·E 2 配合生成故事绘本
2024年12月01日
omegaconf,一个超强的 Python 库!
2024年11月24日
【视觉AIGC识别】误差特征、人脸伪造检测、其他类型假图检测
2024年12月01日
[超级详细]如何在深度学习训练模型过程中使用 GPU 加速
2024年11月29日
Python 物理引擎pymunk最完整教程
2024年11月27日
MediaPipe 人体姿态与手指关键点检测教程
2024年11月27日
深入了解 Taipy:Python 打造 Web 应用的全面教程
2024年11月26日
基于Transformer的时间序列预测模型
2024年11月25日
Python在金融大数据分析中的AI应用(股价分析、量化交易)实战
2024年11月25日
AIGC Gradio系列学习教程之Components
2024年12月01日
Python3 `asyncio` — 异步 I/O,事件循环和并发工具
2024年11月30日
llama-factory SFT系列教程:大模型在自定义数据集 LoRA 训练与部署
2024年12月01日
Python 多线程和多进程用法
2024年11月24日
Python socket详解,全网最全教程
2024年11月27日
python之plot()和subplot()画图
2024年11月26日
理解 DALL·E 2、Stable Diffusion 和 Midjourney 工作原理
2024年12月01日