码迷,mamicode.com
首页 > Web开发 > 详细

web crawling(plus2) comment crawling

时间:2017-10-04 14:14:40      阅读:214      评论:0      收藏:0      [点我收藏+]

标签:mini   urlopen   mozilla   web   data   import   tps   dal   cal   

#Author:Mini
#!/usr/bin/env python
import urllib.request
import re
import urllib.error
headers=("User-Agent","Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:56.0) Gecko/20100101 Firefox/56.0")
opener=urllib.request.build_opener()
opener.addheaders=[headers]
urllib.request.install_opener(opener)
comid="6315179284605088209"
for i in range(4,7):
url="https://coral.qq.com/article/1003390161/comment?commentid="+comid+"&reqnum=20&tag=&callback=jQuery112409214906389806301_1507082215770&_=150708221577"+str(i)

data=urllib.request.urlopen(url).read().decode("utf-8","ignore")
patnext=‘"last":"(.*?)"‘
nextid=re.compile(patnext).findall(data)[0]
patcom=‘"content":"(.*?)",‘
comdata=re.compile(patcom).findall(data)
for j in range(0,len(comdata)):
print(str(i)+"."+str(j)+"comment:")
print(eval(‘u"‘+comdata[j]+‘"‘))
url="https://coral.qq.com/article/1003390161/comment?commentid="+nextid+"&reqnum=20&tag=&callback=jQuery112409214906389806301_1507082215770&_=150708221577"+str(i)
print(url)

web crawling(plus2) comment crawling

标签:mini   urlopen   mozilla   web   data   import   tps   dal   cal   

原文地址:http://www.cnblogs.com/rabbittail/p/7625447.html

(0)
(0)
   
举报
评论 一句话评论(0
登录后才能评论!
© 2014 mamicode.com 版权所有  联系我们:gaon5@hotmail.com
迷上了代码!