【爬虫实战】用python爬今日头条热榜TOP50榜单!
import requests
from bs4 import BeautifulSoup
import pandas as pd
# 设置请求头,模拟浏览器访问
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.3'}
def get_data(url):
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'lxml')
data = soup.find_all('div', class_='title')
rank = [i.span.get_text() for i in soup.find_all('div', class_='num')]
names = [i.a.get_text() for i in data]
hrefs = ['https://www.toutiao.com' + i.a.get('href') for i in data]
return rank, names, hrefs
def main(url):
rank, names, hrefs = get_data(url)
data = pd.DataFrame(list(zip(rank, names, hrefs)), columns=['排名', '名称', '链接'])
print(data)
data.to_csv('今日头条热榜.csv', index=False, encoding='utf-8')
if __name__ == '__main__':
url = 'https://www.toutiao.com/hotwords/'
main(url)
这段代码首先定义了请求头,用于模拟浏览器访问网页。get_data
函数用于获取网页数据,并通过BeautifulSoup进行解析。main
函数则是程序的主要逻辑,它调用get_data
函数获取数据,并将数据存储在一个DataFrame中,最后将数据保存到CSV文件中。最后,在__name__
为__main__
时,执行主函数,开始爬取数据。
评论已关闭