python网络爬虫

时间：2016-10-25 14:06:46 阅读：174 评论：0 收藏：0 [点我收藏+]

标签：超时 link cts text 文件名 ... binary charset family

1.可以通过`r.headers`来获取响应头内容

>>>r = requests.get(‘http://www.zhidaow.com‘)
>>> r.headers
{
    ‘content-encoding‘: ‘gzip‘,
    ‘transfer-encoding‘: ‘chunked‘,
    ‘content-type‘: ‘text/html; charset=utf-8‘;
    ...
}
>>> r.headers[‘Content-Type‘]
‘text/html; charset=utf-8‘

>>> r.headers.get(‘content-type‘)

‘text/html; charset=utf-8‘

2.设置超时时间

>>> requests.get(‘http://github.com‘, timeout=0.001)

3.代理访问

proxies = {
  "http": "http://10.10.1.10:3128",
  "https": "http://10.10.1.10:1080",
}

requests.get("http://www.zhidaow.com", proxies=proxies)

如果代理需要账户和密码，则需这样：

proxies = {
    "http": "http://user:pass@10.10.1.10:3128/",
}

4.不允许重定向的设置

>>>r = requests.get(‘http://www.baidu.com/link?url=QeTRFOS7TuUQRppa0wlTJJr6FfIYI1DJprJukx4Qy0XnsDO_s9baoO8u1wvjxgqN‘, allow_redirects = False)

5. 上传文件

>>> url = ‘http://httpbin.org/post‘
>>> files = {‘file‘: open(‘report.xls‘, ‘rb‘)}

>>> r = requests.post(url, files=files)
>>> r.text
{
  ...
  "files": {
    "file": "<censored...binary...data>"
  },
  ...
}

还可以显式地设置文件名:

>>> url = ‘http://httpbin.org/post‘
>>> files = {‘file‘: (‘report.xls‘, open(‘report.xls‘, ‘rb‘))}

>>> r = requests.post(url, files=files)
>>> r.text
{
  ...
  "files": {
    "file": "<censored...binary...data>"
  },
  ...
}

如果你想，你也可以发送作为文件来接收的字符串:

>>> url = ‘http://httpbin.org/post‘
>>> files = {‘file‘: (‘report.csv‘, ‘some,data,to,send\nanother,row,to,send\n‘)}

>>> r = requests.post(url, files=files)
>>> r.text
{
  ...
  "files": {
    "file": "some,data,to,send\\nanother,row,to,send\\n"
  },
  ...
}

python网络爬虫

标签：超时 link cts text 文件名 ... binary charset family

原文地址：http://www.cnblogs.com/auto-tester-space/p/5996010.html

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行

python网络爬虫

1.可以通过r.headers来获取响应头内容

>>> requests.get(‘http://github.com‘, timeout=0.001)

1.可以通过`r.headers`来获取响应头内容