检索规则说明:AND代表“并且”;OR代表“或者”;NOT代表“不包含”;(注意必须大写,运算符两边需空一格)
检 索 范 例 :范例一: (K=图书馆学 OR K=情报学) AND A=范并思 范例二:J=计算机应用与软件 AND (U=C++ OR U=Basic) NOT M=Visual
机构地区:[1]同济大学计算机科学与工程系,上海200092
出 处:《计算机工程》2003年第2期151-152,160,共3页Computer Engineering
摘 要:汉语自动词性标注和韵律短语切分都是汉语文语转换(Text-to-Speech)系统的重要组成部分。在用从人工标注的语料库中得到韵律短语切分点的边界模式以及概率信息,对文本中的韵律短语切分点进行自动预测时,语素'g'这种词性就过于模糊,导致韵律短语切分点预测得不合理。该文提出了一种修改词类标注集,去掉语素'g'这种词性的方法。该方法在进行词性标注时,对实语素恰当地标注出在句中的词性,以便提高韵律短语的正确切分。应用此方法对10万词的训练集和5万词的测试集分别进行封闭和开放测试表明,词性标注正确率分别可达96.67%和92.60%。并采用修改过的词类标注集,对1000句的文本进行了韵律短语切分点的预测,召回率在66.21%左右,正确率达到了75.79%。Both the Chinese part-of-speech automatic tagging and prosodic phrase segmentation are important modulars in a Chinese text-to-speech system. When predicting phrase breaks using the boundary pattern and boundary distribution probabilities derived from hand-annotated corpus, the authors find that the POS tag 'g' is too ambiguous, which leads to the illogicality of the prediction of phrase breaks. This paper proposes an approach of modifying the POS tag set, so the POS tag 'g' will never be in this set. When tagging part-of-speech for Chinese, in order to improve the correct segmentation of prosodic phrase, the authors annotate morphemes with appropriate POS tags. According to this method train it on a close corpus of 100,000 characters and then test on an open test set of 50,000 characters. The primary experiment proves that the overall accuracy for POS tagging of close corpus and open test set is 96.67% and 92.60% respectively. The authors also test the prediction of phrase breaks on about 1000 sentences using the modified POS tag set, the recalling rate is around 66.21% , the correct rate is about 75.79%.
关 键 词:韵律短语 切分方法 词性标注 词类标注集 语素 汉语信息处理 汉语文语转换系统
分 类 号:TP391.12[自动化与计算机技术—计算机应用技术]
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在链接到云南高校图书馆文献保障联盟下载...
云南高校图书馆联盟文献共享服务平台 版权所有©
您的IP:216.73.216.236