标签:正则表达式笔记
1.所谓组,即你查一次,可以查多种匹配的结果,每种结果即是一组
eg.
import re pattern = re.compile(r'<div(.*?>)(.*?)</div>') match = pattern.match('<div class="test">Hello <div>wa haha<div>test xxx</div></div>world</div>') if match: print match.group(2) print match.group(1) print match.group(0) >>> ================================ RESTART ================================ >>> Hello <div>wa haha<div>test xxx class="test"> <div class="test">Hello <div>wa haha<div>test xxx</div> >>>其中group(0)比较特殊,所有组的合集
2.查找一个文件中所有的http(url)地址
import re import urllib2 f = open('url.txt','r') string = "" while 1: line = f.readline() if not line:break string += line f.close() urls = re.findall('http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\(\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+', string)
标签:正则表达式笔记
原文地址:http://blog.csdn.net/hellochenlian/article/details/44429711