python利用urllib實現的爬取京東網站商品圖片的爬蟲

本文轉載自查看原文 2017-08-23 16:31 1314 python爬蟲

本例程使用urlib實現的，基於python2.7版本，采用beautifulsoup進行網頁分析，沒有第三方庫的應該安裝上之后才能運行，我用的IDE是pycharm，閑話少說，直接上代碼！

 1 # -*- coding: utf-8 -*
 2 import re
 3 import os
 4 import urllib
 5 import urllib2
 6 from bs4 import BeautifulSoup
 7 def craw(url,page):
 8     html1=urllib2.urlopen(url).read()
 9     html1=str(html1)
10     soup=BeautifulSoup(html1,'lxml')
11     imagelist=soup.select('#J_goodsList > ul > li > div > div.p-img > a > img')
12     namelist=soup.select('#J_goodsList > ul > li > div > div.p-name > a > em')
13     #pricelist=soup.select('#plist > ul > li > div > div.p-price > strong')
14     #print pricelist
15     path = "E:/{}/".format(str(goods))
16     if not os.path.exists(path):
17         os.mkdir(path)
18     for (imageurl,name) in zip(imagelist,namelist):
19         name=name.get_text()
20         imagename=path + name  +".jpg"
21         imgurl="http:"+str(imageurl.get('data-lazy-img'))
22         if imgurl == 'http:None':
23             imgurl = "http:" + str(imageurl.get('src'))
24         try:
25             urllib.urlretrieve(imgurl,filename=imagename)
26         except:
27             continue
28 
29 '''
30 #J_goodsList > ul > li:nth-child(1) > div > div.p-img > a > img
31 #plist > ul > li:nth-child(1) > div > div.p-name.p-name-type3 > a > em
32 #plist > ul > li:nth-child(1) > div > div.p-price > strong:nth-child(1) > i
33 '''
34 
35 if __name__ == "__main__":
36     goods=raw_input('please input the goos you want:')
37     pages=input('please input the pages you want:')
38     count =0.0
39     for i in range(1,int(pages+1),2):
40         url="https://search.jd.com/Search?keyword={}&enc=utf-8&qrst=1&rt=1&stop=1&vt=2&suggest=1.def.0.T06&wq=diann&page={}".format(str(goods),str(i))
41         craw(url,i)
42         count += 1
43         print 'work completed {:.2f}%'.format(count/int(pages)*100)

圖片的命名為商品的名稱，京東商品圖片地址的屬性很可能會有所變動，所以大家進行編寫的時候應該舉一反三，靈活運用！
這是我下載下來的手機類圖片文件的截圖：
這里寫圖片描述
我本地的爬取的速度很快，不到一分鍾就能爬取100頁上千個商品的圖片！

免責聲明！

本站轉載的文章為個人學習借鑒使用，本站對版權不負任何法律責任。如果侵犯了您的隱私權益，請聯系本站郵箱yoyou2525@163.com刪除。

猜您在找 利用Python爬蟲爬取京東商品的簡要信息 python爬蟲實踐——爬取京東商品信息網絡爬蟲之網站圖片爬取-python實現 Python爬取京東商品列表 Java爬蟲爬取京東商品信息爬蟲之selenium爬取京東商品信息爬蟲系列(十三) 用selenium爬取京東商品 python爬蟲學習-爬取某個網站上的所有圖片畢設二:python 爬取京東的商品評論 Java爬蟲爬取京東