码迷,mamicode.com
首页 > Web开发 > 详细

web crawling(plus6) pic mining

时间:2017-10-03 16:28:20      阅读:208      评论:0      收藏:0      [点我收藏+]

标签:spi   curl   arc   .com   int   sea   key   firefox   pytho   

#Author:Mini
#!/usr/bin/env python
import urllib.request
import re
import urllib.error
headers=("User-Agent","Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:56.0) Gecko/20100101 Firefox/56.0")
opener=urllib.request.build_opener()
opener.addheaders=[headers]
urllib.request.install_opener(opener)
keyword="连衣裙"
key=urllib.request.quote(keyword)
for i in range(1,101):
try:
url="https://s.taobao.com/search?q="+key+"&imgfile=&js=1&stats_click=search_radio_all%3A1&initiative_id=staobaoz_20171003&ie=utf8&bcoffset=4&ntoffset=4&p4ppushleft=1%2C48&s="+str(i*44)
data=urllib.request.urlopen(url).read().decode("utf-8","ignore")
pat1=‘"pic_url":"//(.*?)"‘
pic=re.compile(pat1).findall(data)
print("success!")
print(pic)
for j in range(0,len(pic)):
thispic=pic[j]
thispicurl="http://"+thispic
picf="E:/m/"+str(i)+"."+str(j)+".jpg"
urllib.request.urlretrieve(thispicurl,filename=picf)
except urllib.error.URLError as e:
if hasattr(e, "code"):
print(e.code)
if hasattr(e, "reason"):
print(e.reason)

web crawling(plus6) pic mining

标签:spi   curl   arc   .com   int   sea   key   firefox   pytho   

原文地址:http://www.cnblogs.com/rabbittail/p/7623819.html

(0)
(0)
   
举报
评论 一句话评论(0
登录后才能评论!
© 2014 mamicode.com 版权所有  联系我们:gaon5@hotmail.com
迷上了代码!