统计英文文档频率前n单词

时间：2017-07-14 19:36:47 阅读：150 评论：0 收藏：0 [点我收藏+]

标签：notepad comm join com 英文 class common 注意需要

#coding:utf-8
#!/usr/bin/python2.6


def statistic_eng_text():
    ‘‘‘统计出英文文档中高频词汇‘‘‘
    cnt = Counter()
    np = os.path.join(get_project_path(),‘doc‘,‘jack lodon.txt‘)
    ff = open(np,‘r‘)
    words = ff.read()
    format_text = re.split(‘[\s\ \\,\;\.\!\n]+‘,words)

    for w in format_text:#比较的时候注意了大小写，其中有一个 the是以大写字母开始的，所以在notepad中统计出来了，而在代码中没有统计出来

        cnt[w.lower()] += 1#这里需要把单词进行一个转换，避免大小写导致的不匹配
    print cnt.most_common(5)

if __name__ == ‘__main__‘:
    statistic_eng_text()

统计英文文档频率前n单词

标签：notepad comm join com 英文 class common 注意需要

原文地址：http://www.cnblogs.com/yNds/p/7171870.html

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行