基于两步策略的中文短文本分类研究  被引量:7

Chinese short-text classification in two-steps

在线阅读下载全文

作  者:樊兴华[1] 王鹏[1] 

机构地区:[1]重庆邮电大学计算机科学与技术研究所,重庆400065

出  处:《大连海事大学学报》2008年第3期121-124,共4页Journal of Dalian Maritime University

基  金:国家自然科学基金资助项目(60703010);重庆市自然科学基金资助项目(2006BB2374);重庆市教委科学技术研究项目(KJ070519);教育部回国留学人员启动基金资助项目(教外司留[2007]1109号)

摘  要:为更好地挖掘文本信息,研究了将两步策略用于中文短文本分类的3个关键问题,提出了基于组合朴素贝叶斯(NB)和K近邻(KNN)分类器的两步中文短文本分类方法:(1)直接利用NB和KNN的输出构造其对应的二维空间,根据该空间内错误文本的分布将测试文本集分为3部分:能被KNN可靠分类的文本集A,不能被KNN可靠分类但能被NB可靠分类的文本集B,其他文本集C.(2)用KNN、NB分别对文本集A和B进行分类,根据训练语料的类别分布,直接给属于文本集C的文本分配标签.与NB、KNN和支持向量机(SVM)的对比实验表明,该方法可获得较高的分类性能.Three key issues of classifying Chinese short-text in two-steps were discussed to mine text information effectively, and a method of combining naive Bayesian (NB) with k-nearest neighbor (KNN) classifiers for this task was developed. Firstly, the test text collection was divided into three parts: part-A which could be classified reliably by KNN, part-B which could not be classified reliably by KNN but could be classified reliably by NB and the another part-C. All above was implemented by utilizing the outputs of NB or KNN classifier to construct the corresponding two-dimension space respectively, and thereby making the division according to the distribution of texts misclassified in the space. Then, part-A and part-B was classified respectively by using KNN and NB classifiers, and part-C was assigned directly the labels according to the distribution of categorization in the training data. The experimental results show that the proposed method achieves high performance comparing with KNN, NB and support vector machine (SVM).

关 键 词:中文短文本 文本分类 两步策略 朴素贝叶斯(NB) K近邻(KNN) 

分 类 号:TP18[自动化与计算机技术—控制理论与控制工程]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象