xpath的使用：定位，获取文本和属性值

时间：2018-10-09 19:03:42 阅读：336 评论：0 收藏：0 [点我收藏+]

标签：hello attr www world str 文档 get classname ext2

myPage = ‘‘‘<html>
<title>TITLE</title>
<body>
<h1></h1>
<div></div>
<div id="photos">
<img src="pic1.jpeg"/><span id="pic1">*</span>
<img src="pic2.jpeg"/><span id="pic2">****</span>
<p><a href="http://www.example.com/more_pic.html">;*</a></p>
<a href="http://www.baidu.com">****</a>;
<a href="http://www.163.com">*****</a>;
<a href="http://www.sohu.com">****</a>;
</div>
<p class="myclassname">Hello,\nworld!<br/>-- by Adam</p>
<div class="foot">放在尾部的其他一些说明</div>
</body>
</html>‘‘‘

html = etree.fromstring(myPage)

#一、定位
divs1 = html.xpath(‘//div‘)
divs2 = html.xpath(‘//div[@id]‘)
divs3 = html.xpath(‘//div[@class="foot"]‘)
divs4 = html.xpath(‘//div[@]‘)
divs5 = html.xpath(‘//div[1]‘)
divs6 = html.xpath(‘//div[last()-1]‘)
divs7 = html.xpath(‘//div[position()<3]‘)
divs8 = html.xpath(‘//div|//h1‘)
divs9 = html.xpath(‘//div[not(@)]‘)

二、取文本 text() 区别 html.xpath(‘string()‘)

text1 = html.xpath(‘//div/text()‘)
text2 = html.xpath(‘//div[@id]/text()‘)
text3 = html.xpath(‘//div[@class="foot"]/text()‘)
text4 = html.xpath(‘//div[@*]/text()‘)
text5 = html.xpath(‘//div[1]/text()‘)
text6 = html.xpath(‘//div[last()-1]/text()‘)
text7 = html.xpath(‘//div[position()<3]/text()‘)
text8 = html.xpath(‘//div/text()|//h1/text()‘)

#三、取属性 @
value1 = html.xpath(‘//a/@href‘)
value2 = html.xpath(‘//img/@src‘)
value3 = html.xpath(‘//div[2]/span/@id‘)

#四、定位（进阶）
#1.文档(DOM)元素(Element)的find，findall方法
divs = html.xpath(‘//div[position()<3]‘)
for div in divs:
ass = div.findall(‘a‘) # 这里只能找到:div->a, 找不到:div->p->a
for a in ass:
if a is not None:
#print(dir(a))
print(a.text, a.attrib.get(‘href‘)) #文档(DOM)元素(Element)的属性：text, attrib

2.与1等价

a_href = html.xpath(‘//div[position()<3]/a/@href‘)
print(a_href)

#3.注意与1、2的区别
a_href = html.xpath(‘//div[position()<3]//a/@href‘)
print(a_href)

参考：https://www.cnblogs.com/hhh5460/p/5079465.html

xpath的使用：定位，获取文本和属性值

标签：hello attr www world str 文档 get classname ext2

原文地址：http://blog.51cto.com/13831593/2296394

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行