一种从XML数据中发现关系信息的方法  被引量:10

A Method of Discovering Relation Information from XML Data

在线阅读下载全文

作  者:吴扬扬[1] 雷庆[1] 陈锻生[1] YOKOTA Harou 

机构地区:[1]华侨大学计算机科学系,福建泉州362021 [2]Department of Computer Science,Tokyo Institute of Technology,Tokyo,Japan

出  处:《软件学报》2008年第6期1422-1427,共6页Journal of Software

基  金:Supported by the Natural Science Foundation of Fujian Province of China under Grant No.A0510020(福建省自然科学基金);the Int'I Science and Technology Cooperation Project of Fujian Province of China under Grant No.20041014(福建省国际科技合作项目)

摘  要:提出了一种发现蕴藏在不同XML文档嵌套结构中的关系信息及其出现模式的新方法.可根据用户兴趣,发现描述不同实体之间联系的关系信息,抽取关系实例及其在文档中的出现模式.具体解决方案是:首先识别和收集包含用户感兴趣的实体的XML文档片段:然后根据文档片段标签的语义和文档片段的结构计算文档片段的相似度,并采用自适应阈值方法按相似度聚类文档片段.使得包含同一种关系的文档片段聚集在同一个片段簇:最后从XML文档片段簇中抽取关系实例及其出现模式.实验结果表明,对于包含有意义标签的各种XML文档,该方法能够准确地识别和抽取出描述指定实体之间联系的各种关系信息.A novel method of discovering relation information among entities buried in different nest structures of XML documents is proposed. The method is able to identify relations among different types of entities given by users, and extract relation instances and their occurrence patterns in XML documents. The solution is as follows: identify and collect XML fragments that contain all types of entity given by users at first, then calculate similarity between fragments based on semantics of their tags and their structures, and cluster fragments with a adaptively selected similarity threshold so that the fragments containing the same relation are clustered together, finally extract relation instances and patterns of their occurrences from each cluster. The experimental results show that the method can identify and extract relation information among given types of entities correctly from all kinds of XML documents with meaningful tags.

关 键 词:关系信息 XML文档 相似度 聚类 出现模式 

分 类 号:TP311[自动化与计算机技术—计算机软件与理论]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象