检索规则说明:AND代表“并且”;OR代表“或者”;NOT代表“不包含”;(注意必须大写,运算符两边需空一格)
检 索 范 例 :范例一: (K=图书馆学 OR K=情报学) AND A=范并思 范例二:J=计算机应用与软件 AND (U=C++ OR U=Basic) NOT M=Visual
机构地区:[1]辽宁石油化工大学计算机与通信工程学院,辽宁抚顺113001
出 处:《辽宁石油化工大学学报》2006年第2期83-86,共4页Journal of Liaoning Petrochemical University
摘 要:互联网上信息量的激增,迫切需要一些自动化的工具帮助人们在海量信息源中迅速找到真正需要的信息,如标题、链接e、mail和图片等,而HTML语言所表述的Web页面经浏览器分析后只适合浏览,不适合作为一种数据交换的方式由机器处理。介绍了HTMLParser的原理和java正则表达式相关知识,基于HTMLParser包和正则表达式。以提取网站内部email信息为例,提出了Web信息抽取系统设计方案,阐述了email信息抽取的工作原理和关键技术,给出了email抽取算法,并详细介绍了系统的抽取URL、email和存储模块,抽取结果保存于数据库中,供机器检索利用。The rapid growth of the Web contents increasers the need for some automatic tools to help to find the exact information among the magnanimous information sources such as titles, links, emails, pictures etc. The Web pages expressed by HTML, after analyzed by Internet Explorer, are suitable for browse, but not for machine processing as the way of data exchange. The principle of HTMLParser and related knowledge of regular expression, package HTMLParser and regular expression were introduced. Taking extracting email information inside websites as an example, the scheme of design was proposed. The principle of email extraction and key technique were presented. The algorithm of email extraction was given. URL extraction module, email extraction module and storage module were described in detail. The result of extraction is stored in database for the use of data retrieval.
关 键 词:信息抽取 正则表达式 HTMLParser包 JAVA
分 类 号:TP311.1[自动化与计算机技术—计算机软件与理论]
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在链接到云南高校图书馆文献保障联盟下载...
云南高校图书馆联盟文献共享服务平台 版权所有©
您的IP:216.73.216.112