python爬取网页数据

本文转载自查看原文 2019-05-08 21:51 1512 python

import re
from urllib.request import urlopen
'''
爬取网页数据信息
'''
def getPage(url):
    response = urlopen(url)
    return response.read().decode('utf-8')

def parsePage(s):
    ret = re.findall(
        '<div class="item">.*?<div class="pic">.*?<em .*?>(?P<id>\d+).*?<span class="title">(?P<title>.*?)</span>'
       '.*?<span class="rating_num" .*?>(?P<rating_num>.*?)</span>.*?<span>(?P<comment_num>.*?)评价</span>',s,re.S)
    return ret

def main(num):
    url = 'https://movie.douban.com/top250?start=%s&filter=' % num
    response_html = getPage(url)
    ret = parsePage(response_html)
    print(ret)

count = 0
for i in range(10):   # 10页
    main(count)
    count += 25

# url从网页上把代码搞下来
# bytes decode ——> utf-8 网页内容就是我的待匹配字符串
# ret = re.findall(正则，带匹配的字符串)  #ret是所有匹配到的内容组成的列表

免责声明！

本站转载的文章为个人学习借鉴使用，本站对版权不负任何法律责任。如果侵犯了您的隐私权益，请联系本站邮箱yoyou2525@163.com删除。

猜您在找 python爬取网页数据方法 python爬取网页数据 python之爬取网页数据总结（一） python爬取网页数据 python爬虫——爬取网页数据和解析数据 python爬虫——爬取网页数据和解析数据 python爬取动态网页数据，详解 Python：将爬取的网页数据写入Excel文件中如何轻松爬取网页数据？ pycharm爬取网页数据