188宝金博页面版

  • 图案背景
  • 纯色背景
视图
标记
批注
批注本地保存成功,开通会员云端永久保存 去开通
adsuhviusa

上传于:2018-09-01

粉丝量:104

该文档贡献者很忙,什么也没留下。


  • 相关
  • 目录
  • 笔记
  • 书签

188宝金博页面版:更多相关文档

  • 基于deep Q-network双足机器人非平整地面行走稳定性控制方法

    星级: 5 页

  • The Deep

    星级: 45 页

  • the deep zoo

    星级: 32 页

  • the deep 01

    星级: 65 页

  • the deep a

    星级: 2 页

  • Deep Q-Network Using Reward Distribution

    星级: 10 页

  • Deep Trouble for the Deep Self…

    星级: 19 页

  • Implementing the

    星级: 1 页

  • Roll in the deep

    星级: 3 页

  • Roll In The Deep

    星级: 3 页

  • Light in the deep sea

    星级: 19 页

  • Roll in the deep

    星级: 4 页

暂无目录

点击鼠标右键菜单,创建目录

暂无笔记

选择文本,点击鼠标右键菜单,添加笔记

暂无书签

在左侧文档中,点击鼠标右键,添加书签

188宝金博页面版: Implementing the Deep Q-Network

下载积分: 300

内容提示: Implementing the Deep Q-NetworkMelrose RoderickHumans To Robots LaboratoryBrown UniversityProvidence, RI 02912melrose_roderick@brown.eduJames MacGlashanHumans To Robots LaboratoryBrown UniversityProvidence, RI 02912jmacglashan@gmail.comStefanie TellexHumans To Robots LaboratoryBrown UniversityProvidence, RI 02912stefie10@cs.brown.eduAbstractThe Deep Q-Network proposed by Mnih et al. [2015] has become a benchmarkand building point for much deep reinforcement learning research. However,replicating results fo...

文档格式:PDF | 页数:9 | 浏览次数:9 | 上传日期:2018-09-01 00:56:51 | 文档星级:
Implementing the Deep Q-NetworkMelrose RoderickHumans To Robots LaboratoryBrown UniversityProvidence, RI 02912melrose_roderick@brown.eduJames MacGlashanHumans To Robots LaboratoryBrown UniversityProvidence, RI 02912jmacglashan@gmail.comStefanie TellexHumans To Robots LaboratoryBrown UniversityProvidence, RI 02912stefie10@cs.brown.eduAbstractThe Deep Q-Network proposed by Mnih et al. [2015] has become a benchmarkand building point for much deep reinforcement learning research. However,replicating results for complex systems is often challenging since original scientif i cpublications are not always able to describe in detail every important parametersetting and software engineering solution. In this paper, we present results fromour work reproducing the results of the DQN paper. We highlight key areas inthe implementation that were not covered in great detail in the original paperto make it easier for researchers to replicate these results, including terminationconditions and gradient descent algorithms. Finally, we discuss methods forimproving the computational performance and provide our own implementationthat is designed to work with a range of domains, and not just the original ArcadeLearning Environment [Bellemare et al., 2013].1 IntroductionOver the past few years, deep reinforcement learning has gained much popularity as it has been shownto perform better than previous methods on domains with very large state-spaces. In one of the earliestdeep reinforcement learning papers (hereafter the DQN paper), Mnih et al. [2015] presented a methodfor learning to play Atari 2600 video games, using the Arcade Learning Environment (ALE) [Belle-mare et al., 2013], from image and performance data alone using the same deep neural networkarchitecture and hyper-parameters for all the games. DQN outperformed previous reinforcementlearning methods on nearly all of the games and recorded better than human performance on most. Asmany researchers tackle reinforcement learning problems with deep reinforcement learning methodsand propose alternative algorithms, the results of the DQN paper are often used as a benchmark toshow improvement. Thus, implementing the DQN algorithm is important for both replicating theresults of the DQN paper for comparison and also building off the original algorithm. One of themain contributions of the DQN paper was f i nding ways to improve stability in their artif i cial neuralnetworks during training. There are, however, a number of other areas in the implementation of thismethod that are crucial to its success, which were only mentioned brief l y in the paper.We implemented a Deep Q-Network (DQN) to play the Atari games and replicated the results ofMnih et al. [2015]. Our implementation, available freely online, 1 runs around 4x faster than the1 www.github.com/h2r/burlap_caffe30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain.arXiv:1711.07478v1 [cs.LG] 20 Nov 2017

188宝金博页面版:关注我们

  • 新浪微博

关注188宝金博页面版公众号

188宝金博页面版
阅读
APP
阅读
返回
顶部
188宝金博页面版官网登录在线平台入口(2026已更新)—江苏协昌电子科技股份有限公司