Query Expansion Using Wikipedia and a Concept Base in Cross-language Information Retrieval  

Query Expansion Using Wikipedia and a Concept Base in Cross-language Information Retrieval

在线阅读下载全文

作  者:Pham Huy Anh Yukawa Takashi 

机构地区:[1]Department o fin formation Science and Technology, Nagaoka University of Technology, Nagaoka-shi 940-2188, Japan

出  处:《Computer Technology and Application》2013年第10期522-531,共10页计算机技术与应用(英文版)

摘  要:The present paper describes the use of online free language resources for translating and expanding queries in CLIR (cross-language information retrieval). In a previous study, we proposed method queries that were translated by two machine translation systems on the Language Gridem. The queries were then expanded using an online dictionary to translate compound words or word phrases. A concept base was used to compare back translation words with the original query in order to delete mistranslated words. In order to evaluate the proposed method, we constructed a CLIR system and used the science documents of the NTCIR1 dataset. The proposed method achieved high precision. However~ proper nouns (names of people and places) appear infrequently in science documents. In information retrieval, proper nouns present unique problems. Since proper nouns are usually unknown words, they are difficult to find in monolingual dictionaries, not to mention bilingual dictionaries. Furthermore, the initial query of the user is not always the best description of the desired information. In order to solve this problem, and to create a better query representation, query expansion is often proposed as a solution. Wikipedia was used to translate compound words or word phrases. It was also used to expand queries together with a concept base. The NTCIRI and NTCIR 6 datasets were used to evaluate the proposed method. In the proposed method, the CLIR system was implemented with a high rate of precision. The proposed syst had a higher ranking than the NTCIRI and NTCIR6 participation systems.

关 键 词:Cross-language inlbrmation retrieval CLIR language resources concept base language grid Wikipedia. 

分 类 号:TP311[自动化与计算机技术—计算机软件与理论] G354.4[自动化与计算机技术—计算机科学与技术]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象