码迷,mamicode.com
首页 > 编程语言 > 详细

python beautifulsoup获取特定html源码

时间:2017-05-12 01:37:42      阅读:208      评论:0      收藏:0      [点我收藏+]

标签:属性   open   url   from   web   blog   print   prettify   pil   

beautifulsoup 获取特定html源码

import re
from bs4 import BeautifulSoup
import urllib2

url = ‘http://www.cnblogs.com/vickey-wu/‘
# connect to a URL
web = urllib2.urlopen(url)
# read html code
html = web.read()
# print html
soup = BeautifulSoup(html,‘html.parser‘)
prety = soup.prettify()
# print prety
pointed_div = soup.findAll(name="div", attrs={"class":re.compile("forFlow")})    # 筛选标签为div且属性class为forFlow的源码
print pointed_div

python beautifulsoup获取特定html源码

标签:属性   open   url   from   web   blog   print   prettify   pil   

原文地址:http://www.cnblogs.com/vickey-wu/p/6843411.html

(0)
(0)
   
举报
评论 一句话评论(0
登录后才能评论!
© 2014 mamicode.com 版权所有  联系我们:gaon5@hotmail.com
迷上了代码!