节点频度和语义距离相结合的网页正文信息抽取被引量：3

Combing node frequency and semantic feature for webpage informative content extraction

出　　处：《计算机工程与应用》2009年第1期140-143,共4页Computer Engineering and Applications

基　　金：国家自然科学基金~~

摘　　要：提出了一种带有节点频度的扩展DOM树模型—BF-DOM树模型(Block node Frequency-Document Object Module),并基于此模型进行网页正文信息的抽取。该方法通过向DOM树的某些节点上添加频度和相关度属性来构造文中新的模型,再结合语义距离抽取网页正文信息。方法主要基于以下三点考虑:在同源的网页集合内噪音节点的频度值很高;正文信息一般由非链接文字组成;与正文相关的链接和文章标题有较近的语义距离。针对8个网站的实验表明,该方法能有效地抽取正文信息,召回率和准确率都在96%以上,优于基于信息熵的抽取方法。A new module named BF-DOM tree is proposed in this paper,which extends the Document Object Module Tree by adding two properties,i.e. ,block node frequency and relativity,to some nodes.Using this module combined with semantic distance, this method extracts the primary content accurately from the same source based on three facts：noise nodes always have high node frequency property within a given website;primary content blocks are often made up of few link words and many text words;useful links are contained in a useful content blocks and have a close semantic distance with page titles.Experiment on eight respective websites shows the proposed method can identify the primary content blocks with higher precision and recall rate both above 96% which is better than the entropy based method.The method can reduce the storage requirement for search engines;thus,result in smaller indexes,faster search time, and better user satisfaction.

关键词：信息提取带有节点频度的文档对象模型树节点频度语义距离

分类号：TP391[自动化与计算机技术—计算机应用技术]

参考文献：

正在载入数据...

二级参考文献：

正在载入数据...

耦合文献：

正在载入数据...

引证文献：

正在载入数据...

二级引证文献：

正在载入数据...

同被引文献：

正在载入数据...

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

节点频度和语义距离相结合的网页正文信息抽取被引量：3

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

高级检索检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

节点频度和语义距离相结合的网页正文信息抽取 被引量：3

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

用户登录

高级检索检索式检索

节点频度和语义距离相结合的网页正文信息抽取被引量：3