python爬虫beautifulsoup4系列3

时间：2017-06-03 12:38:33 阅读：195 评论：0 收藏：0 [点我收藏+]

标签：路径 -o out mvp icon qmvc bcf npm nop

前言

本篇手把手教大家如何爬取网站上的图片，并保存到本地电脑

一、目标网站

1.随便打开一个风景图的网站：http://699pic.com/sousuo-218808-13-1.html

2.用firebug定位，打开firepath里css定位目标图片

3.从下图可以看出，所有的图片都是img标签，class属性都是lazy

技术分享

二、用find_all找出所有的标签

1.find_all(class_="lazy")获取所有的图片对象标签

2.从标签里面提出jpg的url地址和title

 1 # coding:utf-8
 2 from bs4 import BeautifulSoup
 3 import requests
 4 import os
 5 r = requests.get("http://699pic.com/sousuo-218808-13-1.html")
 6 fengjing = r.content
 7 soup = BeautifulSoup(fengjing, "html.parser")
 8 # 找出所有的标签
 9 images = soup.find_all(class_="lazy")
10 # print images # 返回list对象
11 
12 for i in images:
13     jpg_rl = i["data-original"]  # 获取url地址
14     title = i["title"]           # 返回title名称
15     print title
16     print jpg_rl
17     print ""

三、保存图片

1.在当前脚本文件夹下创建一个jpg的子文件夹

2.导入os模块，os.getcwd()这个方法可以获取当前脚本的路径

3.用open打开写入本地电脑的文件路径，命名为：os.getcwd()+"\\jpg\\"+title+‘.jpg‘（命名重复的话，会被覆盖掉）

4.requests里get打开图片的url地址，content方法返回的是二进制流文件，可以直接写到本地

技术分享

四、参考代码

 1 # coding:utf-8
 2 from bs4 import BeautifulSoup
 3 import requests
 4 import os
 5 r = requests.get("http://699pic.com/sousuo-218808-13-1.html")
 6 fengjing = r.content
 7 soup = BeautifulSoup(fengjing, "html.parser")
 8 # 找出所有的标签
 9 images = soup.find_all(class_="lazy")
10 # print images # 返回list对象
11 
12 for i in images:
13     jpg_rl = i["data-original"]
14     title = i["title"]
15     print title
16     print jpg_rl
17     print ""
18     with open(os.getcwd()+"\\jpg\\"+title+‘.jpg‘, "wb") as f:
19         f.write(requests.get(jpg_rl).content)

对python接口自动化有兴趣的，可以加python接口自动化QQ群：226296743

也可以关注下我的个人公众号：

技术分享

python爬虫beautifulsoup4系列3

标签：路径 -o out mvp icon qmvc bcf npm nop

原文地址：http://www.cnblogs.com/yoyoketang/p/6910741.html

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行