Python中，关于读取文件编码解码的问题

时间：2018-11-28 13:13:57 阅读：214 评论：0 收藏：0 [点我收藏+]

标签：def decode code 分享字节 text depend file this

UnicodeDecodeError: ‘gbk‘ codec can‘t decode byte 0xb1 in position 94: illegal multibyte sequence

            有时候用open()方法打开文件读取文件的时候会出现这个问题：‘GBK’编×××无法解码94号位置的字节0xb1：非法多字节序列。错误信息提示了使用“GBK”解码。
            1.分析
            pycharm自动使用的是‘UTF-8’编码，好像没有什么问题，为什么会出现这个错误呢。结果查了下open()函数的注解，里面又这么一段话：
             encoding is the name of the encoding used to decode or encode the  file. This should only be used in text mode. *The default encoding is platform dependent*, but any encoding supported by Python can be  passed.  See the codecs module for the list of supported encodings.
                 The default encoding is platform dependent：默认编码方式取决于平台。这也就不奇怪会用‘GBK’编码了，平台不一样，编码方式不一样，所以读取的时候回出现错误。
            2.解决方法
                    # 1.以byte读取，并以‘utf-8’解码
                    # fp = open(filename, ‘rb‘)
                    # content = fp.read()
                    # self.content = content.decode(‘utf-8‘)
                    # fp.close()
                    # 2.在打开文件时指定编码方式
                    fp = open(filename, encoding=‘utf-8‘)
                    content = fp.read()
                    self.content = content
                    fp.close()

                    如有不同见解，欢迎分享。

标签：def decode code 分享字节 text depend file this

原文地址：http://blog.51cto.com/14094286/2323006

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年07月29日 (22)
2021年07月28日 (40)
2021年07月27日 (32)
2021年07月26日 (79)
2021年07月23日 (29)
2021年07月22日 (30)
2021年07月21日 (42)
2021年07月20日 (16)
2021年07月19日 (90)
2021年07月16日 (35)

周排行