多Agent MDPs中并行Rollout学习算法

Parallel rollout algorithms for multi-agent MDPs

作　　者：李豹[1]

出　　处：《安徽工程大学学报》2014年第2期75-78,共4页Journal of Anhui Polytechnic University

摘　　要：文章在rollout算法基础上研究了在多Agent MDPs的学习问题.利用神经元动态规划逼近方法来降低其空间复杂度,从而减少算法"维数灾".由于Rollout算法具有很强的内在并行性,文中还分析了并行求解方法.通过多级仓库库存控制的仿真试验,验证了Rollout算法在多Agent学习中的有效性.The paper researches Rollout algorithms （RA） for multi-Agent Markov decision processes （MDPs） in the framework of performance potentials theory. Neuro-dynamic programming （NDP） is used to reduce ＂curse of dimensionality＂ of algorithms, Since to rolout algorithms has a very strong intrinsic parallelism,the parallelization method of RA is employed to reduce the time of running algorithms. Finally,an example of multi-level inventory control by using RA under the supply chain environment is provided. The result shows that rollout algorithms are confirmed to be valid in multi-Agent learning.

关键词：ROLLOUT算法神经元动态规划多AGENT学习性能势并行算法

分类号：TP18[自动化与计算机技术—控制理论与控制工程]

参考文献：

正在载入数据...

二级参考文献：

正在载入数据...

耦合文献：

正在载入数据...

引证文献：

正在载入数据...

二级引证文献：

正在载入数据...

同被引文献：

正在载入数据...

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

多Agent MDPs中并行Rollout学习算法

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

高级检索检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

多Agent MDPs中并行Rollout学习算法

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

用户登录

高级检索检索式检索