非平衡数据集分类研究被引量：5

Research on Imbalanced Dataset Learning Method

出　　处：《计算机技术与发展》2011年第9期39-42,共4页Computer Technology and Development

基　　金：国家自然科学基金资助项目(60903203);福建省教育厅A类科技计划项目(JA08222)

摘　　要：现实世界中存在着非平衡数据集,即数据集中的一类样本数量远大于另一类。而少数类样本的识别通常是人们首要关心的,将少数类样本误分为多数类要比将多数类样本误分为少数类付出更高的代价。传统的机器学习算法可能会产生偏向多数类的结果,因而对于少数类而言,预测的效果会很差。在对目前国内外非平衡数据集研究现状深入分析的基础上,针对非平衡数据集数据复杂度研究和失衡解决方法研究两个方向相对孤立及缺乏系统性的缺陷,提出了一种非平衡数据集整体解决框架,以满足日益迫切的应用需求。A dataset is imbalanced if the classification categories are not approximately equally represented.Often real-world datasets are predominately composed of ＂normal＂ examples with only a small percentage of ＂abnormal＂ or ＂interesting＂ examples.It is also the case that the cost of misclassifying an abnormal（interesting） example as a normal example is often much higher than the cost of the reverse error.Traditional machine learning algorithms may be biased towards the majority class,thus producing poor predictive accuracy over the minority class.Based on the deep analysis on current research about rare cases classification,proposes a learning framework to address the problem of relative isolation of research between data complexity and solution of imbalanced data,and lack of systematic defects to meet the increasingly urgent applications.

关键词：非平衡数据集上采样核学习

分类号：TP18[自动化与计算机技术—控制理论与控制工程]

参考文献：

正在载入数据...

二级参考文献：

正在载入数据...

耦合文献：

正在载入数据...

引证文献：

正在载入数据...

二级引证文献：

正在载入数据...

同被引文献：

正在载入数据...

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

非平衡数据集分类研究被引量：5

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

高级检索检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

非平衡数据集分类研究 被引量：5

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

用户登录

高级检索检索式检索

非平衡数据集分类研究被引量：5