检索规则说明:AND代表“并且”;OR代表“或者”;NOT代表“不包含”;(注意必须大写,运算符两边需空一格)
检 索 范 例 :范例一: (K=图书馆学 OR K=情报学) AND A=范并思 范例二:J=计算机应用与软件 AND (U=C++ OR U=Basic) NOT M=Visual
机构地区:[1]清华大学自动化系,北京100084
出 处:《计算机应用》2006年第8期1894-1897,共4页journal of Computer Applications
摘 要:为了有效地提高不均衡数据集中少数类的分类性能,提出了基于初分类的过抽样算法。首先,对测试集进行初分类,以尽可能多地保留多数类的有用信息;其次,对于被初分类预测为少数类的样本进行再次分类,以有效地提高少数类的分类性能。使用美国加州大学欧文分校的数据集将基于初分类的过抽样算法与合成少数类过抽样算法、欠抽样方法进行了实验比较。结果表明,基于初分类的过抽样算法的少数类与多数类的分类性能都优于其他两种算法。To significantly improve the classification performance of the minority class, an over-sampling algorithm based on preliminary classification was presented. Firstly, preliminary classification was made on the test data in order to save the useful information of the majority class as much as possible, Then the test data that were predicted to belong to minority class were reclassified to improve the classification performance of the minority class. Using the data sets provided by University of California, Irvine, the new algorithm was compared with synthetic minority over-sampling technique and under-sampling method. The experimental results show that the new algorithm performs better than the others in terms of the classification performance of the minority class and majority class.
分 类 号:TP311.13[自动化与计算机技术—计算机软件与理论]
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在链接到云南高校图书馆文献保障联盟下载...
云南高校图书馆联盟文献共享服务平台 版权所有©
您的IP:216.73.216.15