Python学习（2）

时间：2017-04-06 23:51:25 阅读：201 评论：0 收藏：0 [点我收藏+]

标签：python

爬取网页的部分链接

#!/usr/bin/python
#coding = utf8
from urllib.request import urlopen
from bs4 import BeautifulSoup
import re
import random
pages = set()
def getlink(pageurl):
    global pages
    html = urlopen(‘http://www.ftchinese.com‘ + pageurl)
    bs_data = BeautifulSoup(html,‘lxml‘)
#from ipdb import set_trace
#set_trace()
    for link in bs_data.find_all(‘a‘,href = re.compile("^(/m/)")):
        if ‘href‘ in link.attrs:
            if link.attrs[‘href‘] not in pages:
            #我们遇到了新页面
                newpage = link.attrs[‘href‘]
                print(newpage)
                pages.add(newpage)
                getlink(newpage)
getlink("")

Python学习（2）

标签：python

原文地址：http://yanruohan.blog.51cto.com/9740053/1913551

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行