一种基于页面Block的Web信息提取方法被引量：3

A Web Information Extraction Algorithm Based on Web Page

出　　处：《计算机技术与发展》2010年第1期197-200,共4页Computer Technology and Development

基　　金：广西自然科学基金(桂科自0640069)

摘　　要：基于页面结构的信息提取是Web数据挖掘中三大研究领域之一。该研究的关键技术是如何识别Web页面的组织形式,从中挖掘所需要的页面信息。文中基于页面的语义分块(Block)给出一个新的块主题提取算法,与传统的以页面为单位的Web信息提取相比,更符合实际情况,粒度优势明显。该算法针对页面中不同分块的重要性给予不同的权值,依据权值大小取舍页面信息提供给用户。针对该算法进行了模拟实验,从实验结果可以看出该算法具有一定的实用性和有效性。Information extraction based web page structure is one of three web data mining s research fields.Key technology of the research is how to recognize web page s organization form and mine the needed information.Intrduces a new block topic-extracted algorithm based on semantic block.Compared with traditional information extraction based on web page,it is more accordant to the fact and the advantage of granularity is evident.This algorithm gives different block weight values according to the importance of different blocks in a web page. Extract useful information for users according to magnitude of block weight. Simulation experiment was preformed for this algorithm. This algorithm has high practicability and effectiveness.

关键词：语义Block Block权值 Block主题提取 WEB信息挖掘

分类号：TP311[自动化与计算机技术—计算机软件与理论]

参考文献：

正在载入数据...

二级参考文献：

正在载入数据...

耦合文献：

正在载入数据...

引证文献：

正在载入数据...

二级引证文献：

正在载入数据...

同被引文献：

正在载入数据...

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

一种基于页面Block的Web信息提取方法被引量：3

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

高级检索检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

一种基于页面Block的Web信息提取方法 被引量：3

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

用户登录

高级检索检索式检索

一种基于页面Block的Web信息提取方法被引量：3