基于可变网格划分的密度偏差抽样算法  被引量:7

Density biased sampling algorithm based on variable grid division

在线阅读下载全文

作  者:盛开元[1] 钱雪忠[1] 吴秦[1] 

机构地区:[1]江南大学物联网工程学院,江苏无锡214122

出  处:《计算机应用》2013年第9期2419-2422,共4页journal of Computer Applications

基  金:国家自然科学基金资助项目(61103129;61202312);江苏省科技支撑计划项目(BE2009009)

摘  要:简单随机抽样是在分析处理大规模数据集时最常用的数据约简方法,但该方法在处理内部分布不均匀的数据集时容易造成类的丢失。基于固定网格划分的密度偏差抽样算法虽能有效解决该问题,但其速度及效果易受网格划分粒度影响。为此提出了基于可变网格划分的密度偏差抽样算法,根据原始数据集每一维的分布特征确定该维相应的划分粒度,进而构建与原始数据集分布特征一致的网格空间。实验结果表明,在可变网格划分的基础上进行密度偏差抽样,样本质量明显提升,而且相对于基于固定网格划分的密度偏差抽样算法,抽样效率亦有所提高。As the most commonly used method of reducing large-scale datasets, simple random sampling usually causes the loss of some clusters when dealing with unevenly distributed dataset. A density biased sampling algorithm based on grid can solve these defects, but both the efficiency and effect of sampling can be affected by the granularity of grid division. To overcome the shortcoming, a density biased sampling algorithm based on variable grid division was proposed. Every dimension of original dataset was divided according to the corresponding distribution, and the structure of the constructed grid was matched with the distribution of original dataset. The experimental results show that density biased sampling based on variable grid division can achieve higher quality of sample dataset and uses less execution time of sampling compared with the density biased sampling algorithm based on fixed grid division.

关 键 词:密度偏差抽样 可变网格划分 数据挖掘 大规模数据集 聚类 

分 类 号:TP181[自动化与计算机技术—控制理论与控制工程] TP301.6[自动化与计算机技术—控制科学与工程]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象