检索规则说明:AND代表“并且”;OR代表“或者”;NOT代表“不包含”;(注意必须大写,运算符两边需空一格)
检 索 范 例 :范例一: (K=图书馆学 OR K=情报学) AND A=范并思 范例二:J=计算机应用与软件 AND (U=C++ OR U=Basic) NOT M=Visual
机构地区:[1]清华大学计算机科学与技术系,北京100084
出 处:《计算机研究与发展》2016年第6期1249-1262,共14页Journal of Computer Research and Development
基 金:国家自然科学基金项目(61103021);国家"八六三"高技术研究发展计划基金项目(2012AA010901)~~
摘 要:基于图形处理器(graphics processing unit,GPU)加速设备的高性能计算机已经成为目前高性能计算领域的一个重要发展趋势.然而,在当前的GPU设备上开发高效的并行程序仍然是一件非常复杂的事情.针对这一问题,1)总结了影响GPU程序性能的5类关键性能指标;2)采用NVIDIA公司提供的CUPTI底层接口,设计并实现了一套GPU程序性能分析工具集,该工具集可以有效地分析GPU程序的性能行为;3)采用该工具集对著名的GPU评测程序集Rodinia中的17个程序和一个真实应用程序进行了负载特征分析.总结出常见性能瓶颈的典型原因,并给出一些开发高效GPU程序的建议.GPU-based high performance computers have become an important trend in the area of high performance computing.However,developing efficient parallel programs on current GPU devices is very complex because of the complex memory hierarchy and thread hierarchy.To address this problem,we summarize five kinds of key metrics that reflect the performance of programs according to the hardware and software architecture.Then we design and implement a performance analysis tool based on underlying CUPTI interfaces provided by NVIDIA, which can collect key metrics automatically without modifying the source code.The tool can analyze the performance behaviors of GPU programs effectively with very little impact on the execution of programs.Finally,we analyze 17 programs in Rodinia benchmark,which is a famous benchmark for GPU programs,and a real application using our tool.By analyzing the value of key metrics,we find the performance bottlenecks of each program and map the bottlenecks back to source code.These analysis results can be used to guide the optimization of CUDA programs and GPU architecture.Result shows that most bottlenecks come from inefficient memory access,and include unreasonable global memory and shared memory access pattern,and low concurrency for these programs.We summarize the common reasons for typical performance bottlenecks and give some high-level suggestions for developing efficient GPU programs.
关 键 词:图形处理器 负载特征分析 RODINIA 硬件计数器 性能指标
分 类 号:TP338.4[自动化与计算机技术—计算机系统结构]
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在链接到云南高校图书馆文献保障联盟下载...
云南高校图书馆联盟文献共享服务平台 版权所有©
您的IP:216.73.216.118