词干单元和卷积神经网络的哈萨克短文本分类  被引量:1

Kazakh Short Text Classification Based on Stem Unit and Convolutional Neural Network

在线阅读下载全文

作  者:沙尔旦尔·帕尔哈提 米吉提·阿不里米提[1] 艾斯卡尔·艾木都拉[1] SARDAR Parhat;MIJIT Ablimit;ASKAR Hamdulla(College of Information Science and Engineering,Xinjiang University,Urumqi 830046,China)

机构地区:[1]新疆大学信息科学与工程学院,乌鲁木齐830046

出  处:《小型微型计算机系统》2020年第8期1627-1633,共7页Journal of Chinese Computer Systems

基  金:国家自然科学基金项目(61662078,61633013)资助;国家重点研发计划项目(2017YFC0820603)资助。

摘  要:针对哈萨克文本分类中词干提取效率低以及传统框架下特征表示维度高、数据稀疏、分类准确率不高等问题,提出基于哈萨克语形态分析的词干提取方法以及wor2vec_TFIDF融合特征表示和卷积神经网络(CNN)的哈萨克短文本分类方法.首先,根据哈萨克语的词素和语音规则,用词-词素平行训练语料训练高效词干提取模型,并用该模型从网上下载的哈萨克短文本中提取词干.其次,用word2vec算法训练词干向量来分布式地表示文本内容,再用TFIDF算法对其进行加权.最后,用CNN进行文本分类实验,得到95.39%的分类准确率.实验结果表明,稳健词素切分及加权词干向量表示和深度学习方法相比传统机器学习方法更能提高哈萨克短文本分类任务的效率.Aiming at the problems of lowefficiency of stem extraction,high dimension of feature representation,data sparsity and lowaccuracy of classification under the traditional framework in Kazakh text classification,proposes a stem extraction method based on Kazakh morphological analysis,and a text classification method based on word2 vec_TFIDF fusion feature representation and convolutional neural network(CNN).Firstly,according to the morpheme and phonetic rules of Kazakh,the high-efficiency stemming model is trained with the word-morpheme parallel training corpus,and the model is used to extract the stems from the Kazakh short texts dow nloaded from the Internet.Secondly,word2 vec algorithm is used to train stem vectors to represent text contents distributedly,and TFIDF algorithm is used to weight them.Finally,uses CNN to conduct text classification experiments,and obtains 95.39%classification accuracy.The experimental results show that robust morpheme segmentation,and weighted stem vector representation and deep learning methods can improve the efficiency of Kazakh short text classification tasks compared with traditional machine learning methods.

关 键 词:哈萨克语 词干提取 词干向量 文本分类 形态学 

分 类 号:TP391[自动化与计算机技术—计算机应用技术]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象