python爬取網頁數據

本文轉載自查看原文 2019-05-08 21:51 1512 python

import re
from urllib.request import urlopen
'''
爬取網頁數據信息
'''
def getPage(url):
    response = urlopen(url)
    return response.read().decode('utf-8')

def parsePage(s):
    ret = re.findall(
        '<div class="item">.*?<div class="pic">.*?<em .*?>(?P<id>\d+).*?<span class="title">(?P<title>.*?)</span>'
       '.*?<span class="rating_num" .*?>(?P<rating_num>.*?)</span>.*?<span>(?P<comment_num>.*?)評價</span>',s,re.S)
    return ret

def main(num):
    url = 'https://movie.douban.com/top250?start=%s&filter=' % num
    response_html = getPage(url)
    ret = parsePage(response_html)
    print(ret)

count = 0
for i in range(10):   # 10頁
    main(count)
    count += 25

# url從網頁上把代碼搞下來
# bytes decode ——> utf-8 網頁內容就是我的待匹配字符串
# ret = re.findall(正則，帶匹配的字符串)  #ret是所有匹配到的內容組成的列表

免責聲明！

本站轉載的文章為個人學習借鑒使用，本站對版權不負任何法律責任。如果侵犯了您的隱私權益，請聯系本站郵箱yoyou2525@163.com刪除。

猜您在找 python爬取網頁數據方法 python爬取網頁數據 python之爬取網頁數據總結（一） python爬取網頁數據 python爬蟲——爬取網頁數據和解析數據 python爬蟲——爬取網頁數據和解析數據 python爬取動態網頁數據，詳解 Python：將爬取的網頁數據寫入Excel文件中如何輕松爬取網頁數據？ pycharm爬取網頁數據