基于词汇链的多文档自动文摘研究  

Research of multi-document summarization based on lexical chains

在线阅读下载全文

作  者:邓箴[1] 包宏[2] 

机构地区:[1]宁夏大学数学计算机学院,宁夏银川750021 [2]北京科技大学信息工程学院,北京100083

出  处:《计算机与应用化学》2012年第11期1384-1386,共3页Computers and Applied Chemistry

基  金:宁夏大学科学研究基金资助项目(项目编号:ZR1122)

摘  要:提出了一种基于词汇链抽取,文法分析的抽取文本代表词条的多文档摘要生成的方法。通过计算词义相似度构建词汇链,结合词频与位置特征进行文本代表词条成员的选择,将含有词条权值高的句子经过聚类形成多文档文摘句集合,然后进行质心句的抽取和排序,生成多文档文摘。该方法不仅考虑了词汇之间的语义信息,还考虑了词条对文本的代表成度,能够改善文摘句抽取的性能。实验结果表明,与单纯的由关键词确定文摘的方法相比,召回率和准确率都有不少的提高。The paper proposes a method for multiple document summarization based on lexical chains exlTaction, grammar and representative word item. In the method, lexical chains are constructed by calculating the semantic similarity between terms, and selected the representative word item according to word frequency and area. Then, It produced the gather of context sensitive which go though clustering the high value of sentences. At last, the center sentence is extracted from each word item. The experimental results shows that the performance of the system has a notable improvement by considering semantic relationship between term. Compared to the method of summarization based on keyword, which has the high of average precision and recall.

关 键 词:多文档文摘 词汇链 聚类 词条 词义相似度 

分 类 号:TP391[自动化与计算机技术—计算机应用技术]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象