python 自學第二課：使用BeautifulSoup抓取鏈接正則表達式

本文轉載自查看原文 2017-11-16 12:57 2690 Python

python 自學第二課：使用BeautifulSoup抓取鏈接正則表達式

具體的查看BeautifulSoup文檔（根據自己的安裝的版本查看對應文檔）
文檔鏈接https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.zh.html#

#!/usr/bin/env python
# -*- coding:utf-8 -*-
import io  
import sys
from urllib import request
from bs4 import BeautifulSoup
import re
sys.stdout = io.TextIOWrapper(sys.stdout.buffer,encoding='utf8') #改變標准輸出的默認編碼  
resp = request.urlopen("http://news.baidu.com/").read().decode("utf-8")
soup =BeautifulSoup(resp,"html.parser")
listUrls=soup.find_all("a",href=re.compile(".*\/\/news\.baidu.*"))
for url in listUrls:
    print (url["href"])

最后效果：

http://news.baidu.com/view.html
http://news.baidu.com/advanced_news.html
http://news.baidu.com/pianhao.html
http://news.baidu.com/n?bypass=lamp&m=pagesother&v=newsgx
http://news.baidu.com/n?cmd=6&loc=0&name=%B1%B1%BE%A9
http://news.baidu.com/history.html
http://news.baidu.com/newscode.html
http://news.baidu.com/licence.html

免責聲明！

本站轉載的文章為個人學習借鑒使用，本站對版權不負任何法律責任。如果侵犯了您的隱私權益，請聯系本站郵箱yoyou2525@163.com刪除。

猜您在找 c#使用正則表達式抓取a標簽的鏈接和innerhtml Python 正則表達式的使用 python 正則表達式使用 Python正則表達式抓取郵箱分形樹的繪制——python第二課 selenium第二課（腳本錄制seleniumIDE的使用） Delphi 之第二課類的定義與使用 Selenium+python --使用正則表達式爬取頁面的URL鏈接正則表達式抓取文件內容中的http鏈接地址 js項目練習第二課

python 自學第二課： 使用BeautifulSoup抓取鏈接 正則表達式

python 自學第二課： 使用BeautifulSoup抓取鏈接 正則表達式

免責聲明！

python 自學第二課：使用BeautifulSoup抓取鏈接正則表達式

python 自學第二課：使用BeautifulSoup抓取鏈接正則表達式