188宝金博页面版

  • 图案背景
  • 纯色背景
视图
标记
批注
批注本地保存成功,开通会员云端永久保存 去开通
adsuhviusa

上传于:2018-08-29

粉丝量:104

该文档贡献者很忙,什么也没留下。


  • 相关
  • 目录
  • 笔记
  • 书签

188宝金博页面版:更多相关文档

  • lstm进行词性标注

    星级: 18 页

  • 基于Bi-LSTM的医疗事件识别研究

    星级: 7 页

  • LSTM学习介绍,各种门的介绍

    星级: 4 页

  • RNN及LSTM

    星级: 23 页

  • 了解LSTM网络

    星级: 7 页

  • LSTM网络的制造工具111(已完成)

    星级: 5 页

  • 通过多任务知识辅助LSTM识别文本的完成

    星级: 14 页

  • phi-LSTM: A Phrase-Based Hierarchical LSTM Model for Image Captioning

    星级: 17 页

  • 简化LSTM的语音合成

    星级: 5 页

  • 基于LSTM的用户负荷区间预测方法

    星级: 3 页

  • 基于LSTM的对话状态追踪模型研究与实现

    星级: 65 页

暂无目录

点击鼠标右键菜单,创建目录

暂无笔记

选择文本,点击鼠标右键菜单,添加笔记

暂无书签

在左侧文档中,点击鼠标右键,添加书签

188宝金博页面版: Revisiting the Hierarchical Multiscale LSTM

下载积分: 500

内容提示: Revisiting the Hierarchical Multiscale LSTM?Akos K? ad? arTilburg Universitya.kadar@uvt.nlMarc-Alexandre C? ot? eMicrosoft Research Montrealmacote@microsoft.comGrzegorz Chrupa laTilburg Universityg.chrupala@uvt.nlAfra AlishahiTilburg Universitya.alishahi@uvt.nlAbstractHierarchical Multiscale LSTM (Chung et al., 2016a) is a state-of-the-art language modelthat learns interpretable structure from character-level input. Such models can pro-vide fertile ground for (cognitive) computational linguistics stud...

文档格式:PDF | 页数:13 | 浏览次数:15 | 上传日期:2018-08-29 09:46:27 | 文档星级:
Revisiting the Hierarchical Multiscale LSTM´Akos K´ ad´ arTilburg Universitya.kadar@uvt.nlMarc-Alexandre Cˆ ot´ eMicrosoft Research Montrealmacote@microsoft.comGrzegorz Chrupa laTilburg Universityg.chrupala@uvt.nlAfra AlishahiTilburg Universitya.alishahi@uvt.nlAbstractHierarchical Multiscale LSTM (Chung et al., 2016a) is a state-of-the-art language modelthat learns interpretable structure from character-level input. Such models can pro-vide fertile ground for (cognitive) computational linguistics studies. However, the highcomplexity of the architecture, training procedure and implementations might hinder itsapplicability. We provide a detailed reproduction and ablation study of the architecture,shedding light on some of the potential caveats of re-purposing complex deep-learningarchitectures. We further show that simplifying certain aspects of the architecture can infact improve its performance. We also investigate the linguistic units (segments) learnedby various levels of the model, and argue that their quality does not correlate with theoverall performance of the model on language modeling.1 IntroductionVerifying and reproducing claims published in scientif i c articles is an essential part of building asolid foundation for future research. As such, reproduction has a long history in many scientif i cf i elds (Willett et al., 1985; Venables et al., 1993; Waltemath et al., 2011). Recent large-scalestudies, however, raise concerns about the reproducibility in a variety of areas and the potentialef f ect of this crisis (Baker, 2016). A reproduction study by the Open Science Collaboration (2015)estimates that only 40% of research in psychology is reproducible, while Begley and Ellis (2012)end up conf i rming only 11% of preclinical cancer studies. The latter work makes an importantlink between the low reproducibility rates and the notoriously low impact of preclinical cancerresearch on clinical practice (Hutchinson and Kirk, 2011).Our work is motivated by a similar concern, specif i cally the applicability of complex deeplearning architectures for computational (cognitive) linguistics research. State of the art sys-tems often employ complex architectures which integrate various design features and use manyoptimization techniques. Because the focus is on boosting the f i nal performance on a given task,often little ef f ort is put into understanding where the power of the system comes from. Thismakes these models much harder to adapt for new tasks or domains. In our view, it is essentialnot only to be able to reproduce reported results, but also to understand the contribution ofvarious design features through systematic ablation experiments.The higher performance brought by modern neural network architectures often comes at thecost of our understanding of the representations and structural information the system learns.However, for models to be generalizable to new domains, it is important to move towards ana-lyzing such structural representations, and investigating their impact on the f i nal performanceof the model.In the current study, we examine the reproducibility of a language model with the abilityto learn explicit linguistic structure: the Hierarchical Multiscale Recurrent Neural Network(HMLSTM) model. This architecture was introduced by Chung et al. (2016a) and set a newstate of the art on language-modeling benchmarks Text8, Hutter Prize and character-level PennTreebank. Additionally, the paper features examples where the lowest layer of the model recoversarXiv:1807.03595v1 [cs.CL] 10 Jul 2018

188宝金博页面版:关注我们

  • 新浪微博

关注188宝金博页面版公众号

188宝金博页面版
阅读
APP
阅读
返回
顶部
188宝金博页面版官网登录在线平台入口(2026已更新)—江苏协昌电子科技股份有限公司