Normalization of Homophonic Words in Chinese Microblogs  

在线阅读下载全文

作  者:Xin Zhang Jiaying Song Yu He Guohong Fu 

机构地区:[1]School of Computer Science and Technology,Heilongjiang University Harbin150080,China

出  处:《国际计算机前沿大会会议论文集》2015年第1期51-53,共3页International Conference of Pioneering Computer Scientists, Engineers and Educators(ICPCSEE)

基  金:This study was supported by National Natural Science Foundation of China under Grant No.61170148 and No.60973081, the Returned Scholar Foundation of Heilongjiang Province, and Harbin Innovative Foundation for Returnees under Grant No.2009RFLXG007, respectively.

摘  要:Homophonic words are very popular in Chinese microblog, posing a new challenge for Chinese microblog text analysis. However, to date, there has been very little research conducted on Chinese homophonic words normalization. In this paper, we take Chinese homophonic word normalization as a process of language decoding and propose an n-gram based approach. To this end, we first employ homophonic–original word or character mapping tables to generate normalization candidates for a given sentence with homophonic words, and thus exploit n-gram language models to decode the best normalization from the candidate set. Our experimental results show that using the homophonic-original character mapping table and n-grams trained from the microblog corpus help improve performance in homophonic word recognition and restoration.

关 键 词:Microblog analysis TEXT NORMALIZATION homophonic WORDS N-GRAM 

分 类 号:C5[社会学]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象