中文科技文献切分的领域适应技术研究

Research on Domain Adaptation Technology of Chinese Science and Technology Literatures Segmentation

出　　处：《图书情报工作》2014年第19期13-18,共6页Library and Information Service

基　　金：科技部国际科技合作专项"面向科技文献的日汉双向实用型机器翻译合作研究"(项目编号:2014DFA11350);国家社会科学基金项目"基于事实型科技大数据的情报分析方法及集成分析平台研究"(项目编号:14BTQ038)研究成果之一

摘　　要：以生物医学文献为实例对象,研究科技文献切分中的领域适应技术,通过以词典特征、领域词汇特征、子串标注和使用词典切分的粗切分语料作为训练语料等方法,实现基于序列标注的中文切分方法由新闻领域到科技领域的适应,并取得了较好的效果。研究表明,在科技文献切分中,充分利用领域知识获取领域相关特征,对于提高科技文献切分的准确率具有重要的作用。Segmentation of science and technology（S＆T） literature is a basic step in S＆T documents information processing. This paper takes biomedical literatures as the instances and studies domain adaptation technology in segmentation of S＆T literatures. Then it takes some methods such as dictionary features, domain character features, sub-word tagging and low quality in-domain training corpus based on dictionary-based segmentation to adapt Chinese segmentation method based on sequence labeling in journalism filed to S＆T filed and achieves the significant improvement. It finds that how to exploit domain specific features with domain knowledge plays an important role in improving the segmentation quality of S＆T literatures.

关键词：中文切分领域适应科技文献信息处理

分类号：TP391.1[自动化与计算机技术—计算机应用技术]

参考文献：

正在载入数据...

二级参考文献：

正在载入数据...

耦合文献：

正在载入数据...

引证文献：

正在载入数据...

二级引证文献：

正在载入数据...

同被引文献：

正在载入数据...

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

中文科技文献切分的领域适应技术研究

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

高级检索检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

中文科技文献切分的领域适应技术研究

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

用户登录

高级检索检索式检索