使用Urllib爬虫(1)--简单的将数据爬到内存或硬盘中

时间：2020-04-18 10:05:54 阅读：91 评论：0 收藏：0 [点我收藏+]

将数据爬取到内存中

import urllib
import urllib.request
import re
#打开京东网页并且进行读取，解码格式utf-8,ignore小细节自动略过，大大减少出错率
#将数据爬到内存中
#http://www.jd.com
url = "http://www.jd.com"
data = urllib.request.urlopen(url).read().decode("utf-8","ignore")
pat = "<title>(.*?)</title>"
#re.S模式修正符，网页数据往往是多行的，避免多行的影响
print(re.compile(pat,re.S).findall(data))

将数据爬取到硬盘中

import urllib
import urllib.request
import re
url = "http://www.jd.com"
#urlretrieve(网址，文件名filename),由于\有转义的作用所以改用为/或者\\
res = urllib.request.urlretrieve(url,filename="D:\\pythonstudy\\pachong\\jd1.html")
print(res)

使用Urllib爬虫(1)--简单的将数据爬到内存或硬盘中

标签：http -- 爬取 code 作用网页数据 ons 解码细节

原文地址：https://www.cnblogs.com/u-damowang1/p/12724139.html

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行