提高韵律短语正确切分方法的研究  

Research on the Approach of Improving the Correct Segmentation of Prosodic Phrase

在线阅读下载全文

作  者:吴晓慧[1] 柴佩琪[1] 

机构地区:[1]同济大学计算机科学与工程系,上海200092

出  处:《计算机工程》2003年第2期151-152,160,共3页Computer Engineering

摘  要:汉语自动词性标注和韵律短语切分都是汉语文语转换(Text-to-Speech)系统的重要组成部分。在用从人工标注的语料库中得到韵律短语切分点的边界模式以及概率信息,对文本中的韵律短语切分点进行自动预测时,语素'g'这种词性就过于模糊,导致韵律短语切分点预测得不合理。该文提出了一种修改词类标注集,去掉语素'g'这种词性的方法。该方法在进行词性标注时,对实语素恰当地标注出在句中的词性,以便提高韵律短语的正确切分。应用此方法对10万词的训练集和5万词的测试集分别进行封闭和开放测试表明,词性标注正确率分别可达96.67%和92.60%。并采用修改过的词类标注集,对1000句的文本进行了韵律短语切分点的预测,召回率在66.21%左右,正确率达到了75.79%。Both the Chinese part-of-speech automatic tagging and prosodic phrase segmentation are important modulars in a Chinese text-to-speech system. When predicting phrase breaks using the boundary pattern and boundary distribution probabilities derived from hand-annotated corpus, the authors find that the POS tag 'g' is too ambiguous, which leads to the illogicality of the prediction of phrase breaks. This paper proposes an approach of modifying the POS tag set, so the POS tag 'g' will never be in this set. When tagging part-of-speech for Chinese, in order to improve the correct segmentation of prosodic phrase, the authors annotate morphemes with appropriate POS tags. According to this method train it on a close corpus of 100,000 characters and then test on an open test set of 50,000 characters. The primary experiment proves that the overall accuracy for POS tagging of close corpus and open test set is 96.67% and 92.60% respectively. The authors also test the prediction of phrase breaks on about 1000 sentences using the modified POS tag set, the recalling rate is around 66.21% , the correct rate is about 75.79%.

关 键 词:韵律短语 切分方法 词性标注 词类标注集 语素 汉语信息处理 汉语文语转换系统 

分 类 号:TP391.12[自动化与计算机技术—计算机应用技术]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象