python_爬虫_str类型的html文本去标签

本文转载自查看原文 2018-09-05 17:20 1323 Python_爬虫

# from HTMLParser import HTMLParser
from html.parser import HTMLParser # 将字符串格式的html文本转成html

class MyHTMLParser(HTMLParser):
    def __init__(self):
        HTMLParser.__init__(self)
        self.data = []
    def handle_startendtag(self, tag, attrs):
        pass
    def handle_endtag(self, tag):
        pass
    def handle_data(self, data):
        if data.count('\n') == 0:
            self.data.append(data)

if __name__ == '__main__':
    parser = MyHTMLParser()
    for i in conn(): # 获取文章
        content = i[0]
        parser.feed(content)

        parser.data # 通过这个可以获取去标签后的内容列表

参考：https://www.cnblogs.com/AlwinXu/p/5492033.html

免责声明！

本站转载的文章为个人学习借鉴使用，本站对版权不负任何法律责任。如果侵犯了您的隐私权益，请联系本站邮箱yoyou2525@163.com删除。

猜您在找 python 正则提取HTml标签文本内容的 Python_对象类型判断 python之路--str类型 Python_网络爬虫（新浪新闻抓取） Python_报错：TypeError: write() argument must be str, not int python_爬虫_multiprocessing.dummy以及multiprocessing Python_异常：TypeError: write() argument must be str, not list HTML 文本标签数据爬虫：使用python爬取HTML标签 python之str基础类型