'# Python实时爬虫:自动抓取并推送学校最新通知
一、背景与问题
在校园信息化建设中,通知公告的实时获取是提升办公效率的关键环节。传统手动查看通知的方式存在以下痛点:
- 时效性差:人工检查需要持续关注页面更新
- 信息遗漏:多平台通知分散管理容易漏看
- 操作成本高:每日重复性检查消耗人力
- 数据沉淀难:缺乏结构化存储导致信息难以追溯
本项目通过构建实时爬虫系统,实现以下核心价值:
- 自动抓取指定网页的最新通知
- 实时推送至指定渠道(如手机通知)
- 历史通知自动归档并支持快速检索
二、基本原理
系统架构包含四个核心模块:
- 爬虫采集模块:使用requests/selenium获取网页内容
- 数据解析模块:使用BeautifulSoup/PyQuery提取结构化数据
- 数据存储模块:使用SQLite/MySQL存储历史数据
- 消息推送模块:通过Pushover/Telegram发送实时通知
关键技术点包括:
- 反爬虫策略:处理验证码、IP封禁、请求频率限制
- 动态内容处理:应对JavaScript渲染的网页内容
- 异常处理机制:确保系统健壮性
- 定时任务调度:使用schedule库实现定时爬取
三、环境准备
pip install requests beautifulsoup4 selenium schedule pushover-api需要准备的环境要素:
- 浏览器驱动:ChromeDriver(用于处理动态网页)
- Pushover账户:注册获取API token
- 数据库配置:SQLite或MySQL的连接信息
- 代理服务:应对IP封禁时的代理配置
四、核心实现
1. 爬虫采集模块
import requests
from bs4 import BeautifulSoup
import time
def fetch_page(url, headers):
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
return response.text
except Exception as e:
print(f"请求失败: {e}")
return None关键点解释:
- 使用headers参数模拟浏览器访问
- 添加超时控制防止卡死
- 异常处理保证程序稳定性
2. 动态内容处理(Selenium示例)
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
def get_dynamic_content(url):
chrome_options = Options()
chrome_options.add_argument("--headless") # 无头模式
driver = webdriver.Chrome(options=chrome_options)
try:
driver.get(url)
time.sleep(3) # 等待动态内容加载
return driver.page_source
finally:
driver.quit()注意事项:
- 使用无头模式避免浏览器界面弹出
- 需要处理动态加载的延迟问题
- 可结合selenium-wire进行请求监控
3. 数据解析模块
def parse_notice(html):
soup = BeautifulSoup(html, 'html.parser')
notices = []
for item in soup.select('.notice-item'):
title = item.select_one('.title').text.strip()
date = item.select_one('.date').text.strip()
link = item.select_one('a')['href']
notices.append({
'title': title,
'date': date,
'link': link
})
return notices优化建议:
- 使用CSS选择器提高解析效率
- 对异常数据进行清洗处理
- 可扩展支持多种网页结构
五、完整案例
1. 学校通知爬虫完整流程
import sqlite3
from pushover import Client
# 配置信息
DB_NAME = 'notices.db'
PUSHOVER_TOKEN = 'your_token'
PUSHOVER_USER = 'your_user'
def init_db():
conn = sqlite3.connect(DB_NAME)
c = conn.cursor()
c.execute('''CREATE TABLE IF NOT EXISTS notices
(id INTEGER PRIMARY KEY, title TEXT, date TEXT, link TEXT, timestamp DATETIME)''')
conn.commit()
conn.close()
def main():
url = 'https://example.edu/notice'
headers = {
'User-Agent': 'Mozilla/5.0',
'Referer': 'https://example.edu'
}
html = get_dynamic_content(url)
if not html:
return
notices = parse_notice(html)
# 历史数据对比
conn = sqlite3.connect(DB_NAME)
c = conn.cursor()
c.execute("SELECT MAX(timestamp) FROM notices")
last_time = c.fetchone()[0]
new_notices = []
for notice in notices:
if not last_time or notice['timestamp'] > last_time:
new_notices.append(notice)
# 存储新数据
c.executemany("INSERT INTO notices (title, date, link, timestamp) VALUES (?, ?, ?, ?)",
[(n['title'], n['date'], n['link'], datetime.now()) for n in new_notices])
conn.commit()
conn.close()
# 推送新通知
client = Client(PUSHOVER_USER, token=PUSHOVER_TOKEN)
for notice in new_notices:
client.push(title=notice['title'], message=f"新通知:{notice['date']}\n{notice['link']}")完整流程说明:
- 使用Selenium获取动态加载的页面内容
- 解析提取通知标题、日期、链接
- 对比历史记录发现新增通知
- 将新通知存入SQLite数据库
- 通过Pushover推送至指定设备
六、源码解析
1. 动态内容处理机制
def get_dynamic_content(url):
chrome_options = Options()
chrome_options.add_argument("--headless")
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("--no-sandbox")
chrome_options.add_argument(f"--proxy-server=http:{PROXY_SERVER}")
driver = webdriver.Chrome(options=chrome_options)
try:
driver.get(url)
time.sleep(5) # 等待动态内容加载
return driver.page_source
finally:
driver.quit()关键点:
- 无头模式避免浏览器界面弹出
- 代理配置防止IP被封
- 等待时间需根据页面加载速度调整
2. 数据存储优化
def batch_insert(notices):
conn = sqlite3.connect(DB_NAME)
c = conn.cursor()
c.executemany("INSERT OR IGNORE INTO notices (title, date, link, timestamp) VALUES (?, ?, ?, ?)",
[(n['title'], n['date'], n['link'], datetime.now()) for n in notices])
conn.commit()
conn.close()优化策略:
- 使用
INSERT OR IGNORE避免重复插入 - 批量操作提高效率
- 可扩展为MySQL的批量插入
七、进阶使用
1. 分布式爬虫架构
from multiprocessing import Pool
def process_page(url):
html = get_dynamic_content(url)
if html:
return parse_notice(html)
return []
def distributed_crawler(urls):
with Pool(processes=4) as p:
results = p.map(process_page, urls)
return [item for sublist in results for item in sublist]适用场景:
- 多源数据采集需求
- 需要并行处理多个网页
- 服务器资源充足时
2. 性能优化策略
from functools import lru_cache
@lru_cache(maxsize=100)
def get_cached_page(url):
return fetch_page(url, headers)优化方向:
- 使用缓存减少重复请求
- 实现请求队列管理
- 使用数据库存储访问记录
八、性能与工程实践
1. 并发控制
from threading import Semaphore
MAX_CONCURRENCY = 5
semaphore = Semaphore(MAX_CONCURRENCY)
def safe_fetch(url):
with semaphore:
return fetch_page(url, headers)注意事项:
- 避免对服务器造成过大压力
- 设置合理的并发数
- 可结合速率限制策略
2. 异常处理机制
def safe_request(url):
try:
return fetch_page(url, headers)
except requests.exceptions.RequestException as e:
print(f"请求异常: {e}")
return None
except Exception as e:
print(f"未知异常: {e}")
return None处理策略:
- 区分不同类型的异常
- 设置重试机制
- 记录错误日志
九、常见问题与踩坑
1. 反爬虫机制应对
错误示例:
def fetch_page(url):
return requests.get(url).text问题分析:
- 缺少User-Agent
- 未处理验证码
- 未设置请求间隔
改进方案:
def fetch_page(url):
headers = {
'User-Agent': 'Mozilla/5.0',
'Referer': 'https://example.edu'
}
return requests.get(url, headers=headers, timeout=10).text2. 推送服务配置问题
常见错误:
- 未正确设置API token
- 未处理推送失败情况
解决办法:
def send_pushover(title, message):
client = Client(PUSHOVER_USER, token=PUSHOVER_TOKEN)
try:
response = client.push(title=title, message=message)
if response.status_code != 200:
print("推送失败")
except Exception as e:
print(f"推送异常: {e}")十、最佳实践
1. 系统架构建议
├── config/ # 配置文件
├── logs/ # 日志文件
├── src/ # 源代码
│ ├── crawler/ # 爬虫模块
│ ├── parser/ # 解析模块
│ ├── storage/ # 存储模块
│ └── notifier/ # 推送模块
├── db/ # 数据库
└── requirements.txt # 依赖文件2. 安全实践建议
- 使用HTTPS加密通信
- 对敏感信息进行加密存储
- 设置访问权限控制
- 定期更换API密钥
十一、总结
本项目通过构建完整的爬虫系统,实现了学校通知的自动抓取与推送。关键点包括:
- 动态内容处理:使用Selenium应对JavaScript渲染的网页
- 异常处理机制:确保系统在异常情况下稳定运行
- 数据存储优化:使用SQLite进行结构化存储
- 消息推送服务:通过Pushover实现即时通知
适用场景:
- 需要实时获取特定网页数据
- 有多个信息源需要整合
- 需要自动化处理数据的场景
不适用场景:
- 非法爬取受保护数据
- 需要处理大量复杂数据结构
- 对数据精度要求极高的场景
通过合理的设计和优化,本系统可以稳定运行于生产环境,为用户提供及时的信息服务。在实施过程中需注意法律风险和安全防护,确保系统在合规的前提下运行。