188宝金博页面版

  • 图案背景
  • 纯色背景
视图
标记
批注
批注本地保存成功,开通会员云端永久保存 去开通
adsuhviusa

上传于:2018-08-31

粉丝量:104

该文档贡献者很忙,什么也没留下。


  • 相关
  • 目录
  • 笔记
  • 书签

188宝金博页面版:更多相关文档

  • ML-Deep Neural Networks

    星级: 18 页

  • with deep neural networks

    星级: 6 页

  • Deep Neural Networks

    星级: 21 页

  • Deep Neural Networks

    星级: 75 页

  • A Case for Neural Networks

    星级: 7 页

  • Learn Keras for Deep Neural Networks

    星级: 234 页

  • Deep Neural Networks—A Brief History

    星级: 18 页

  • DeepFault Fault Localization for Deep Neural Networks

    星级: 21 页

  • DyVEDeep Dynamic Variable Effort Deep Neural Networks

    星级: 16 页

  • Neural Networks Library in Java

    星级: 5 页

  • Deep Q-Networks for Accelerating the Training of Deep Neural Networks

    星级: 10 页

  • A representer theorem for deep neural networks.pdf

    星级: 22 页

  • A Novel Diagnosis Method for SZ by Deep Neural Networks

    星级: 9 页

  • A Learning Technique for Deep Belief Neural Networks

    星级: 11 页

  • Deep and Modular Neural Networks

    星级: 22 页

暂无目录

点击鼠标右键菜单,创建目录

暂无笔记

选择文本,点击鼠标右键菜单,添加笔记

暂无书签

在左侧文档中,点击鼠标右键,添加书签

188宝金博页面版: DeepThin: A Self-Compressing Library for Deep Neural Networks

下载积分: 500

内容提示: DeepThin: A Self-Compressing Library forDeep Neural NetworksMatthew Sotoudeh ?Intel/UC Davismasotoudeh@ucdavis.eduSara S. BaghsorkhiIntelsara.s.baghsorkhi@intel.comAbstractAs the industry deploys increasingly large and complexneural networks to mobile devices, more pressure isput on the memory and compute resources of those de-vices.Deepcompression,orcompressionofdeepneuralnetwork weight matrices, is a technique to stretch re-sources for such scenarios. Existing compression meth-ods cannot ef f ectively c...

文档格式:PDF | 页数:15 | 浏览次数:12 | 上传日期:2018-08-31 13:53:26 | 文档星级:
DeepThin: A Self-Compressing Library forDeep Neural NetworksMatthew Sotoudeh ∗Intel/UC Davismasotoudeh@ucdavis.eduSara S. BaghsorkhiIntelsara.s.baghsorkhi@intel.comAbstractAs the industry deploys increasingly large and complexneural networks to mobile devices, more pressure isput on the memory and compute resources of those de-vices.Deepcompression,orcompressionofdeepneuralnetwork weight matrices, is a technique to stretch re-sources for such scenarios. Existing compression meth-ods cannot ef f ectively compress models smaller than1-2% of their original size. We develop a new compres-siontechnique,DeepThin,buildingonexistingresearchin the area of low rank factorization. We identify andbreak artif i cial constraints imposed by low rank ap-proximations by combining rank factorization with areshaping process that adds nonlinearity to the approx-imation function.We deploy DeepThin as a plug-gablelibraryintegratedwithTensorFlowthatenablesuserstoseamlessly compress models at dif f erent granularities.We evaluate DeepThin on two state-of-the-art acous-tic models, TFKaldi and DeepSpeech, comparing it toprevious compression work (Pruning, HashNet, andRank Factorization), empirical limit study approaches,and hand-tuned models. For TFKaldi, our DeepThinnetworks show better word error rates (WER) thancompeting methods at practically all tested compres-sion rates, achieving an average of 60% relative im-provement over rank factorization, 57% over pruning,23% over hand-tuned same-size networks, and 6% overthe computationally expensive HashedNets. For Deep-Speech, DeepThin-compressed networks achieve bettertest loss than all other compression methods, reachinga 28% better result than rank factorization, 27% betterthan pruning, 20% better than hand-tuned same-sizenetworks, and 12% better than HashedNets.DeepThin also provide inference performance bene-f i ts in two ways: (1) by shrinking the application work-ing sets, allowing the model to f i t in a level of the∗ Work done while at Intel.arXiv, February, 2018cache/memory hierarchy where the original networkwas too large, and (2) by exploiting unique featuresof the technique to reuse many intermediate computa-tions, reducing the total compute operations necessary.We evaluate the performance of DeepThin inferenceacross three Haswell- and Broadwell-based platformswith varying cache sizes. Speedups range from 2X to14X, depending on the compression ratio and platformcache sizes.1 Introduction and MotivationIn recent years, machine learning algorithms have beenincreasingly used in consumer-facing products, suchas speech recognition in personal assistants. These al-gorithms rely on large weight matrices which encodethe relationships between dif f erent nodes in a network.Ideally, these algorithms would run directly on theclient devices such as Amazon Echo [ 20 ] and GoogleHome [ 14 ], which utilize them. Unfortunately, becausesuch devices are usually portable, low-power devices,running such expensive algorithms is infeasible due tothe signif i cant storage, performance, and energy limi-tations involved.To work around this problem, many developers haveresorted to executing the inference models on high-performance cloud servers and streaming the inputsand outputs of the model between client and server.This solution, however, introduces many issues, includ-ing high operational costs, use of large amounts of datatransfer on metered (mobile) networks, user privacyconcerns, and increased latency.Recent research has investigated methods of com-pressing models to sizes which can be ef f i ciently exe-cuted directly on the client device. Such compressionapproaches must reduce the model space requirementwithout signif i cantly impacting prediction accuracy,runtime performance, or engineering time. Our workbuilds on existing research in the area of low rank fac-torization. We develop a new compression method andlibrary, DeepThin, that:arXiv:1802.06944v1 [cs.LG] 20 Feb 2018

188宝金博页面版:关注我们

  • 新浪微博

关注188宝金博页面版公众号

188宝金博页面版
阅读
APP
阅读
返回
顶部
188宝金博页面版官网登录在线平台入口(2026已更新)—江苏协昌电子科技股份有限公司