用shell分析nginx日志百度网页蜘蛛列表页来访情况

时间：2014-12-17 13:01:18 阅读：185 评论：0 收藏：0 [点我收藏+]

#!/bin/bash
#desc: this scripts for baidunews-spider
#date:2014.02.25
#testd in CentOS 5.9 x86_64
#saved in /usr/local/bin/baidu-web.sh
#written by coralzd@gmail.com www.zjyxh.com
dt=`date -d "yesterday" +%m%d`
if [ $1x != x ] ;then
  if [ -e $1 ] ;then
     grep -i "Baiduspider/2.0" $1 > baiduspider-${dt}.txt
     num=`cat baiduspider-${dt}.txt|wc -l`
     echo "baiduspider number is ${num},file is baidu-${dt}.txt"
     cat baiduspider-${dt}.txt|awk ‘{print $7}‘|sort |uniq -c|sort -r >`ls ${1}|cut -c 1-10`-${dt}.txt
     echo "$1 was done"
    else
       echo "$1 not exsist!"
  fi
else
     echo "usage: $0 file_path"
fi

本次用shell分析百度网页蜘蛛跟百度新闻蜘蛛一个方法，无非就是把关键词由baiduspider-news换为baiduspider/2.0。

本文出自 “崔晓辉的博客” 博客，请务必保留此出处http://coralzd.blog.51cto.com/90341/1590956

标签：百度网页 shell baiduspider

原文地址：http://coralzd.blog.51cto.com/90341/1590956

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行