Python实时爬虫:自动抓取并推送学校最新通知

'# Python实时爬虫:自动抓取并推送学校最新通知

一、背景与问题

在校园信息化建设中,通知公告的实时获取是提升办公效率的关键环节。传统手动查看通知的方式存在以下痛点:

  1. 时效性差:人工检查需要持续关注页面更新
  2. 信息遗漏:多平台通知分散管理容易漏看
  3. 操作成本高:每日重复性检查消耗人力
  4. 数据沉淀难:缺乏结构化存储导致信息难以追溯

本项目通过构建实时爬虫系统,实现以下核心价值:

  • 自动抓取指定网页的最新通知
  • 实时推送至指定渠道(如手机通知)
  • 历史通知自动归档并支持快速检索

二、基本原理

系统架构包含四个核心模块:

  1. 爬虫采集模块:使用requests/selenium获取网页内容
  2. 数据解析模块:使用BeautifulSoup/PyQuery提取结构化数据
  3. 数据存储模块:使用SQLite/MySQL存储历史数据
  4. 消息推送模块:通过Pushover/Telegram发送实时通知

关键技术点包括:

  • 反爬虫策略:处理验证码、IP封禁、请求频率限制
  • 动态内容处理:应对JavaScript渲染的网页内容
  • 异常处理机制:确保系统健壮性
  • 定时任务调度:使用schedule库实现定时爬取

三、环境准备

pip install requests beautifulsoup4 selenium schedule pushover-api

需要准备的环境要素:

  1. 浏览器驱动:ChromeDriver(用于处理动态网页)
  2. Pushover账户:注册获取API token
  3. 数据库配置:SQLite或MySQL的连接信息
  4. 代理服务:应对IP封禁时的代理配置

四、核心实现

1. 爬虫采集模块

import requests
from bs4 import BeautifulSoup
import time

def fetch_page(url, headers):
    try:
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        return response.text
    except Exception as e:
        print(f"请求失败: {e}")
        return None

关键点解释:

  • 使用headers参数模拟浏览器访问
  • 添加超时控制防止卡死
  • 异常处理保证程序稳定性

2. 动态内容处理(Selenium示例)

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

def get_dynamic_content(url):
    chrome_options = Options()
    chrome_options.add_argument("--headless")  # 无头模式
    driver = webdriver.Chrome(options=chrome_options)
    try:
        driver.get(url)
        time.sleep(3)  # 等待动态内容加载
        return driver.page_source
    finally:
        driver.quit()

注意事项:

  • 使用无头模式避免浏览器界面弹出
  • 需要处理动态加载的延迟问题
  • 可结合selenium-wire进行请求监控

3. 数据解析模块

def parse_notice(html):
    soup = BeautifulSoup(html, 'html.parser')
    notices = []
    for item in soup.select('.notice-item'):
        title = item.select_one('.title').text.strip()
        date = item.select_one('.date').text.strip()
        link = item.select_one('a')['href']
        notices.append({
            'title': title,
            'date': date,
            'link': link
        })
    return notices

优化建议:

  • 使用CSS选择器提高解析效率
  • 对异常数据进行清洗处理
  • 可扩展支持多种网页结构

五、完整案例

1. 学校通知爬虫完整流程

import sqlite3
from pushover import Client

# 配置信息
DB_NAME = 'notices.db'
PUSHOVER_TOKEN = 'your_token'
PUSHOVER_USER = 'your_user'

def init_db():
    conn = sqlite3.connect(DB_NAME)
    c = conn.cursor()
    c.execute('''CREATE TABLE IF NOT EXISTS notices
                 (id INTEGER PRIMARY KEY, title TEXT, date TEXT, link TEXT, timestamp DATETIME)''')
    conn.commit()
    conn.close()

def main():
    url = 'https://example.edu/notice'
    headers = {
        'User-Agent': 'Mozilla/5.0',
        'Referer': 'https://example.edu'
    }
    
    html = get_dynamic_content(url)
    if not html:
        return
    
    notices = parse_notice(html)
    
    # 历史数据对比
    conn = sqlite3.connect(DB_NAME)
    c = conn.cursor()
    c.execute("SELECT MAX(timestamp) FROM notices")
    last_time = c.fetchone()[0]
    
    new_notices = []
    for notice in notices:
        if not last_time or notice['timestamp'] > last_time:
            new_notices.append(notice)
    
    # 存储新数据
    c.executemany("INSERT INTO notices (title, date, link, timestamp) VALUES (?, ?, ?, ?)",
                  [(n['title'], n['date'], n['link'], datetime.now()) for n in new_notices])
    conn.commit()
    conn.close()
    
    # 推送新通知
    client = Client(PUSHOVER_USER, token=PUSHOVER_TOKEN)
    for notice in new_notices:
        client.push(title=notice['title'], message=f"新通知:{notice['date']}\n{notice['link']}")

完整流程说明:

  1. 使用Selenium获取动态加载的页面内容
  2. 解析提取通知标题、日期、链接
  3. 对比历史记录发现新增通知
  4. 将新通知存入SQLite数据库
  5. 通过Pushover推送至指定设备

六、源码解析

1. 动态内容处理机制

def get_dynamic_content(url):
    chrome_options = Options()
    chrome_options.add_argument("--headless")
    chrome_options.add_argument("--disable-gpu")
    chrome_options.add_argument("--no-sandbox")
    chrome_options.add_argument(f"--proxy-server=http:{PROXY_SERVER}")
    
    driver = webdriver.Chrome(options=chrome_options)
    try:
        driver.get(url)
        time.sleep(5)  # 等待动态内容加载
        return driver.page_source
    finally:
        driver.quit()

关键点:

  • 无头模式避免浏览器界面弹出
  • 代理配置防止IP被封
  • 等待时间需根据页面加载速度调整

2. 数据存储优化

def batch_insert(notices):
    conn = sqlite3.connect(DB_NAME)
    c = conn.cursor()
    c.executemany("INSERT OR IGNORE INTO notices (title, date, link, timestamp) VALUES (?, ?, ?, ?)",
                  [(n['title'], n['date'], n['link'], datetime.now()) for n in notices])
    conn.commit()
    conn.close()

优化策略:

  • 使用INSERT OR IGNORE避免重复插入
  • 批量操作提高效率
  • 可扩展为MySQL的批量插入

七、进阶使用

1. 分布式爬虫架构

from multiprocessing import Pool

def process_page(url):
    html = get_dynamic_content(url)
    if html:
        return parse_notice(html)
    return []

def distributed_crawler(urls):
    with Pool(processes=4) as p:
        results = p.map(process_page, urls)
    return [item for sublist in results for item in sublist]

适用场景:

  • 多源数据采集需求
  • 需要并行处理多个网页
  • 服务器资源充足时

2. 性能优化策略

from functools import lru_cache

@lru_cache(maxsize=100)
def get_cached_page(url):
    return fetch_page(url, headers)

优化方向:

  • 使用缓存减少重复请求
  • 实现请求队列管理
  • 使用数据库存储访问记录

八、性能与工程实践

1. 并发控制

from threading import Semaphore

MAX_CONCURRENCY = 5
semaphore = Semaphore(MAX_CONCURRENCY)

def safe_fetch(url):
    with semaphore:
        return fetch_page(url, headers)

注意事项:

  • 避免对服务器造成过大压力
  • 设置合理的并发数
  • 可结合速率限制策略

2. 异常处理机制

def safe_request(url):
    try:
        return fetch_page(url, headers)
    except requests.exceptions.RequestException as e:
        print(f"请求异常: {e}")
        return None
    except Exception as e:
        print(f"未知异常: {e}")
        return None

处理策略:

  • 区分不同类型的异常
  • 设置重试机制
  • 记录错误日志

九、常见问题与踩坑

1. 反爬虫机制应对

错误示例:

def fetch_page(url):
    return requests.get(url).text

问题分析:

  • 缺少User-Agent
  • 未处理验证码
  • 未设置请求间隔

改进方案:

def fetch_page(url):
    headers = {
        'User-Agent': 'Mozilla/5.0',
        'Referer': 'https://example.edu'
    }
    return requests.get(url, headers=headers, timeout=10).text

2. 推送服务配置问题

常见错误:

  • 未正确设置API token
  • 未处理推送失败情况

解决办法:

def send_pushover(title, message):
    client = Client(PUSHOVER_USER, token=PUSHOVER_TOKEN)
    try:
        response = client.push(title=title, message=message)
        if response.status_code != 200:
            print("推送失败")
    except Exception as e:
        print(f"推送异常: {e}")

十、最佳实践

1. 系统架构建议

├── config/                # 配置文件
├── logs/                 # 日志文件
├── src/                  # 源代码
│   ├── crawler/          # 爬虫模块
│   ├── parser/           # 解析模块
│   ├── storage/          # 存储模块
│   └── notifier/         # 推送模块
├── db/                   # 数据库
└── requirements.txt      # 依赖文件

2. 安全实践建议

  • 使用HTTPS加密通信
  • 对敏感信息进行加密存储
  • 设置访问权限控制
  • 定期更换API密钥

十一、总结

本项目通过构建完整的爬虫系统,实现了学校通知的自动抓取与推送。关键点包括:

  1. 动态内容处理:使用Selenium应对JavaScript渲染的网页
  2. 异常处理机制:确保系统在异常情况下稳定运行
  3. 数据存储优化:使用SQLite进行结构化存储
  4. 消息推送服务:通过Pushover实现即时通知

适用场景:

  • 需要实时获取特定网页数据
  • 有多个信息源需要整合
  • 需要自动化处理数据的场景

不适用场景:

  • 非法爬取受保护数据
  • 需要处理大量复杂数据结构
  • 对数据精度要求极高的场景

通过合理的设计和优化,本系统可以稳定运行于生产环境,为用户提供及时的信息服务。在实施过程中需注意法律风险和安全防护,确保系统在合规的前提下运行。

最后修改于:2026年09月24日 16:00

评论已关闭

推荐阅读

AIGC实战——Transformer模型
2024年12月01日
Socket TCP 和 UDP 编程基础(Python)
2024年11月30日
python , tcp , udp
如何使用 ChatGPT 进行学术润色?你需要这些指令
2024年12月01日
AI
最新 Python 调用 OpenAi 详细教程实现问答、图像合成、图像理解、语音合成、语音识别(详细教程)
2024年11月24日
ChatGPT 和 DALL·E 2 配合生成故事绘本
2024年12月01日
omegaconf,一个超强的 Python 库!
2024年11月24日
【视觉AIGC识别】误差特征、人脸伪造检测、其他类型假图检测
2024年12月01日
[超级详细]如何在深度学习训练模型过程中使用 GPU 加速
2024年11月29日
Python 物理引擎pymunk最完整教程
2024年11月27日
MediaPipe 人体姿态与手指关键点检测教程
2024年11月27日
深入了解 Taipy:Python 打造 Web 应用的全面教程
2024年11月26日
基于Transformer的时间序列预测模型
2024年11月25日
Python在金融大数据分析中的AI应用(股价分析、量化交易)实战
2024年11月25日
AIGC Gradio系列学习教程之Components
2024年12月01日
Python3 `asyncio` — 异步 I/O,事件循环和并发工具
2024年11月30日
llama-factory SFT系列教程:大模型在自定义数据集 LoRA 训练与部署
2024年12月01日
Python 多线程和多进程用法
2024年11月24日
Python socket详解,全网最全教程
2024年11月27日
python之plot()和subplot()画图
2024年11月26日
理解 DALL·E 2、Stable Diffusion 和 Midjourney 工作原理
2024年12月01日