模拟登陆+数据爬取 (python+selenuim)

时间：2017-12-19 01:24:33 阅读：133 评论：0 收藏：0 [点我收藏+]

标签：pytho text tps wrap 为什么 bsp bs4 email position

以下代码是用来爬取LinkedIn网站一些学者的经历的，仅供参考，注意：不要一次性大量爬取会被封号，不要问我为什么知道

#-*- coding:utf-8 -*-
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
import time
from bs4 import BeautifulSoup

diver=webdriver.Chrome()
diver.get(‘https://www.linkedin.com/‘)
#等待网站加载完成
time.sleep(1)
#模拟登陆
diver.find_element_by_id(‘login-email‘).send_keys(用户名)
diver.find_element_by_id(‘login-password‘).send_keys(密码)
# 点击跳转
diver.find_element_by_id(‘login-submit‘).send_keys(Keys.ENTER)
time.sleep(1)
#查询
 diver.find_element_by_tag_name(‘input‘).send_keys(学者名)
diver.find_element_by_tag_name(‘input‘).send_keys(Keys.ENTER)
time.sleep(1)
#获取当前页面所有可能的人
soup=BeautifulSoup(diver.page_source,‘lxml‘)
items=soup.findAll(‘div‘,{‘class‘:‘search-result__wrapper‘})
n=0
for i in items:
n+=1
title=i.find(‘div‘,{‘class‘:‘search-result__image-wrapper‘}).find(‘a‘)[‘href‘]
diver.get(‘https://www.linkedin.com‘+title)
time.sleep(3)
Soup=BeautifulSoup(diver.page_source,‘lxml‘)
# print Soup
Items=Soup.findAll(‘li‘,{‘class‘:‘pv-profile-section__card-item pv-position-entity ember-view‘})
print str(n)+‘:‘
for i in Items:
    print i.find(‘div‘,{‘class‘:‘pv-entity__summary-info‘}).get_text().replace(‘\n‘,‘‘)
diver.close()

标签：pytho text tps wrap 为什么 bsp bs4 email position

原文地址：http://www.cnblogs.com/ybf-yyj/p/8059171.html

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行