Creating customized data services from web pages  

Creating customized data services from web pages

在线阅读下载全文

作  者:季光 Wang Guiling Han Yanbo 

机构地区:[1]Institute of Computing Technology,Chinese Academy of Sciences [2]Graduate University of Chinese Academy of Sciences [3]Research Center for Cloud Computing,North China University of Technology

出  处:《High Technology Letters》2013年第2期203-207,共5页高技术通讯(英文版)

基  金:Supported by the National High Technology Research and Development Programme of China(No.2009AA01 Z141);the National Natural Science Foundation of China(No.60573117);Beijing Natural Science Foundation(No.4131001)

摘  要:To extract structured data from a web page with customized requirements,a user labels some DOM elements on the page with attribute names.The common features of the labeled elements are utilized to guide the user through the labeling process to minimize user efforts,and are also utilized to retrieve attribute values.To turn the attribute values into a structured result,the attribute pattern needs to be induced.For this purpose,a space-optimized suffix tree called attribute tree is built to transform the document object model(DOM) tree into a simpler form while preserving its useful properties such as attribute sequence order.The pattern is induced bottom-up on the attribute tree,and is further used to build the structured result.Experiments are conducted and show high performance of our approach in terms of precision,recall and structural correctness.To extract structured data from a web page with customized requirements, a user labels some DOM elements on the page with attribute names. The common features of the labeled elements are utilized to guide the user through the labeling process to minimize user efforts, and are also utilized to retrieve attribute values. To turn the attribute values into a structured result, the attribute pattern needs to be induced. For this purpose, a space-optimized suffix tree called attribute tree is built to transform the document object model (DOM) tree into a simpler form while preserving its useful properties such as attribute sequence order. The pattern is induced bottom-up on the attribute tree, and is further used to build the structured result. Experiments are conducted and show high perform- ance of our approach in terms of precision, recall and structural correctness.

关 键 词:web data extraction structured data user labeling CUSTOMIZATION data service 

分 类 号:TP393.092[自动化与计算机技术—计算机应用技术] P315.69[自动化与计算机技术—计算机科学与技术]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象