Python+ Calibre 处理中文报纸

import re
##2
line='<a href=nw.D110000renmrb_20180401_1-01.htm><script>document.write(view("领航新时代中国经济航船  "))</script></a>'
#line = line.decode("utf-8")
filtrate = re.compile(u'[^u4E00-u9FA5]')#非中文
filtered_str = filtrate.sub(r'', line)#replace
print (filtered_str)

##1
tempLine = '<script>document.write(view("加强党中央对经济工作的集中统一领导<br>打好决胜全面建成小康社会三大攻坚战  "))</script>'
filtrate = re.compile(u'[^u4E00-u9FA5]')#非中文
filtered_str = filtrate.sub(r'', tempLine)#replace
print (filtered_str)

filtrate = re.findall (r"[u4e00-u9fa5]+", tempLine)

print (filtrate)

python 提取中文

相关阅读:
Python交互设计_接口设计
hibernate注解——@Temporal
java日期格式处理
Unknown tag
个人总结
学习进度条——第十七周
学习进度条——第十六周
学习进度条——第十五周
第二阶段冲刺——个人总结10
学习进度条——十四周

原文地址：https://www.cnblogs.com/xuanyuanchen/p/8721896.html

Python+ Calibre 处理 中文报纸

Python+ Calibre 处理中文报纸