python爬蟲之獲取頁面script里面的內容

本文轉載自查看原文 2019-02-11 19:35 10077

這是網頁上的script 我要獲取的是00914這個數字直接使用正則表達式即可

運行結果：

源碼：

import re
from bs4 import BeautifulSoup
from urllib.request import urlopen
url = "你要解析的網頁URL"
html = urlopen(url).read()
soup = BeautifulSoup(html,"html.parser")
titles = soup.select("body  script") # CSS 選擇器
i = 1
for title in titles:
    if i == 3:
     #print(title.get_text())# 標簽體、標簽屬性
     str=title.get_text()
     break
    if i == 2:
        i = 3
    if i == 1:
        i = 2

print(str)
str1 = "\"\"\""+"<script>"+str+"</script>"+"\"\"\""
soup = BeautifulSoup(str1, "html.parser")
pattern = re.compile(r"var _url = '(.*?)';$", re.MULTILINE | re.DOTALL)
script = soup.find("script", text=pattern)
#print (pattern.search(script.text).string)
s = pattern.search(script.text).string
print (s.split('\'')[11])

免責聲明！

本站轉載的文章為個人學習借鑒使用，本站對版權不負任何法律責任。如果侵犯了您的隱私權益，請聯系本站郵箱yoyou2525@163.com刪除。

猜您在找 如何獲取HttpServletResponse里面的內容 python獲取script里的內容 python讀取word里面的內容獲取當前頁面的所有鏈接的四種方法對比（python 爬蟲）通過id獲取指定元素內容（標簽里面的標簽內容獲取） PHP獲取微信頁面的指定內容從IE瀏覽器獲取當前頁面的內容 IFrame里面的子頁面html內容變化時，怎么動態改變IFrame的高度？ c#遍歷枚舉里面的內容-獲取枚舉數量，取得大小