python beautifulsoup获取特定html源码

本文转载自查看原文 2017-05-11 23:18 3529 python/ beautifulsoup

beautifulsoup 获取特定html源码（无需登录页面）

import re
from bs4 import BeautifulSoup
import urllib2

url = 'http://www.cnblogs.com/vickey-wu/'
# connect to a URL
web = urllib2.urlopen(url)
# read html code
html = web.read()
# print html
soup = BeautifulSoup(html,'html.parser')
prety = soup.prettify()
# print prety
pointed_div = soup.findAll(name="div", attrs={"class":re.compile("forFlow")})　　　　# 筛选标签为div且属性class为forFlow的源码
print pointed_div

免责声明！

本站转载的文章为个人学习借鉴使用，本站对版权不负任何法律责任。如果侵犯了您的隐私权益，请联系本站邮箱yoyou2525@163.com删除。

猜您在找 【Python】 html解析BeautifulSoup [学习]用python的BeautifulSoup分析html Python 使用 beautifulsoup 4 模块来处理 HTML BeautifulSoup去除html中的标签，获取文本 python用BeautifulSoup解析源码时，去除空格及换行符 Python网页解析：BeautifulSoup vs lxml.html Python3.x的BeautifulSoup解析html常用函数 Python获取网页指定内容(BeautifulSoup工具的使用方法) python 模块BeautifulSoup使用 python bs4 BeautifulSoup