检索规则说明:AND代表“并且”;OR代表“或者”;NOT代表“不包含”;(注意必须大写,运算符两边需空一格)
检 索 范 例 :范例一: (K=图书馆学 OR K=情报学) AND A=范并思 范例二:J=计算机应用与软件 AND (U=C++ OR U=Basic) NOT M=Visual
作 者:李文[1,2] 苗夺谦[1] 卫志华[1] 王炜立[1,2]
机构地区:[1]同济大学计算机科学与技术系,上海201804 [2]南昌大学信息工程学院,南昌330031
出 处:《模式识别与人工智能》2010年第4期456-463,共8页Pattern Recognition and Artificial Intelligence
基 金:国家自然科学基金(No.60475019;60775036;60970061);教育部博士点专项基金(No.20060247039)资助项目
摘 要:文本层次分类中阻塞现象是影响层次分类器性能的重要原因.针对这一问题,提出基于阻塞先验知识的文本层次分类模型.该模型包括两部分:首先对阻塞分布进行估计,提出"阻塞对"识别技术,重点在于获取严重的阻塞方向;其次,把分析出的阻塞先验知识融合到分类过程中,利用层次拓扑结构修正算法,引导阻塞文本"回归"正确分类路径.在中文语料TanCorp上的实验表明,该算法在没有额外增加分类器数目的前提下,能有效改善层次分类性能,是解决层次分类阻塞问题的一种方法.另外,与平面分类算法比较后,该算法更稳定.Blocking exerts negative effect on the performance of text hierarchical classification. In this paper, a two-step hierarchical text classification model based on blocking priori knowledge is proposed to address the problem. Firstly, blocking distribution is estimated and blocking pair recognition technique focusing on mining the serious blocking direction is presented. Secondly, the hierarchy topology structure is actively refined which attempts to correct misclassification and reduce blocking errors by using blocking priori knowledge. The experimental results on TanCorp, which is a new corpus special for Chinese.text classification, show that the model can improve the performance significantly without increasing the extra number of classifiers and is a method of solving the hierarchical classification blocking problem. In addition, compared with fiat text classification algorithm, this method has stable performance.
分 类 号:TP391.1[自动化与计算机技术—计算机应用技术]
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在链接到云南高校图书馆文献保障联盟下载...
云南高校图书馆联盟文献共享服务平台 版权所有©
您的IP:216.73.216.117