基于HTMLParser的Web信息抽取系统的设计与实现被引量：8

Design and Implementation of Web Information Extraction System Based on HTMLParser

机构地区：[1]辽宁石油化工大学计算机与通信工程学院,辽宁抚顺113001

出　　处：《辽宁石油化工大学学报》2006年第2期83-86,共4页Journal of Liaoning Petrochemical University

摘　　要：互联网上信息量的激增,迫切需要一些自动化的工具帮助人们在海量信息源中迅速找到真正需要的信息,如标题、链接e、mail和图片等,而HTML语言所表述的Web页面经浏览器分析后只适合浏览,不适合作为一种数据交换的方式由机器处理。介绍了HTMLParser的原理和java正则表达式相关知识,基于HTMLParser包和正则表达式。以提取网站内部email信息为例,提出了Web信息抽取系统设计方案,阐述了email信息抽取的工作原理和关键技术,给出了email抽取算法,并详细介绍了系统的抽取URL、email和存储模块,抽取结果保存于数据库中,供机器检索利用。The rapid growth of the Web contents increasers the need for some automatic tools to help to find the exact information among the magnanimous information sources such as titles, links, emails, pictures etc. The Web pages expressed by HTML, after analyzed by Internet Explorer, are suitable for browse, but not for machine processing as the way of data exchange. The principle of HTMLParser and related knowledge of regular expression, package HTMLParser and regular expression were introduced. Taking extracting email information inside websites as an example, the scheme of design was proposed. The principle of email extraction and key technique were presented. The algorithm of email extraction was given. URL extraction module, email extraction module and storage module were described in detail. The result of extraction is stored in database for the use of data retrieval.

关键词：信息抽取正则表达式 HTMLParser包 JAVA

分类号：TP311.1[自动化与计算机技术—计算机软件与理论]

参考文献：

正在载入数据...

二级参考文献：

正在载入数据...

耦合文献：

正在载入数据...

引证文献：

正在载入数据...

二级引证文献：

正在载入数据...

同被引文献：

正在载入数据...

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

基于HTMLParser的Web信息抽取系统的设计与实现被引量：8

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

高级检索检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

基于HTMLParser的Web信息抽取系统的设计与实现 被引量：8

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

用户登录

高级检索检索式检索

基于HTMLParser的Web信息抽取系统的设计与实现被引量：8