突发事件热点话题识别系统及关键问题研究  被引量:6

Study on hot topics identification and key issues about emergency events

在线阅读下载全文

作  者:陈莉萍[1] 杜军平[1] 

机构地区:[1]北京邮电大学计算机学院,北京100876

出  处:《计算机工程与应用》2011年第32期19-22,共4页Computer Engineering and Applications

基  金:国家自然科学基金No.91024001;No.61070142;中央高校基本科研业务费专项资金资助(No.2009RC0210);北京市自然科学基金项目(No.4111002)~~

摘  要:针对突发事件热点话题识别系统,建立了系统实现的整体技术框架,给出了系统四个组成部分的关键问题描述及解决策略,结合新闻报道文本内容和结构的特点和报道源分布性特征,基于VSM文本表示模型和TF-IDF公式,提出了正文裁剪方法和特征权重计算的改进模型,并以地震突发事件新闻报道作为数据源进行模型评估。实验结果表明通过对新闻报道正文的裁剪,只提取标题、导语及相关特征参量等信息即可作为热点话题识别的样本集,且改进的特征权重计算模型与经典模型比较,具有更好地执行效率和适应性更强的文本表示能力。Concerning the system of hot topics detection about the emergency events,an overall technical framework is established to implement the system.Description and solution strategy about the key issues in the four components of the system are provided.In terms of the content and structure features of the news reports as well as the distribution feature of the report sources, the text clipping method and the modified model of feature weighting calculation are proposed based on the VSM text representation model and the TF-IDF formula.The news reports about the earthquake emergency event are evaluated for this model as the data sources.Experimental results indicate that the information such as the headline,the lead and relevant feature parameters by clipping the main body of the news report can be considered as the sample set of the hot topics to be identified.Furthermore, compared with the classical model,the modified feature items weighting calculation model is more efficient in execution and more adantive in terms of the text representation capability.

关 键 词:突发事件 新闻报道 热点话题识别 正文裁剪 文本表示模型 

分 类 号:TP391[自动化与计算机技术—计算机应用技术]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象