188宝金博页面版

  • 图案背景
  • 纯色背景
视图
标记
批注
批注本地保存成功,开通会员云端永久保存 去开通
adsuhviusa

上传于:2018-09-01

粉丝量:104

该文档贡献者很忙,什么也没留下。


  • 相关
  • 目录
  • 笔记
  • 书签

188宝金博页面版:更多相关文档

  • 深度学习数学基础 Mathematics of Deep Learning

    星级: 112 页

  • deep learning

    星级: 10 页

  • DEEP LEARNING

    星级: 5 页

  • Deep learning

    星级: 24 页

  • deep learning

    星级: 137 页

  • Deep Learning

    星级: 708 页

  • deep learning

    星级: 7 页

  • deep learning

    星级: 43 页

  • Deep Learning

    星级: 18 页

  • deep learning

    星级: 10 页

  • Deep Learning

    星级: 13 页

  • Deep Learning

    星级: 7 页

  • Deep Learning

    星级: 643 页

  • Deep Learning

    星级: 3 页

  • Deep Learning

    星级: 120 页

暂无目录

点击鼠标右键菜单,创建目录

暂无笔记

选择文本,点击鼠标右键菜单,添加笔记

暂无书签

在左侧文档中,点击鼠标右键,添加书签

188宝金博页面版: Mathematics of Deep Learning

下载积分: 500

内容提示: Mathematics of Deep LearningRené Vidal Joan Bruna Raja Giryes Stefano SoattoAbstract—Recently there has been a dramatic increase in theperformance of recognition systems due to the introduction ofdeep architectures for representation learning and classif i cation.However, the mathematical reasons for this success remainelusive. This tutorial will review recent work that aims toprovide a mathematical justif i cation for several properties ofdeep networks, such as global optimality, geometric stability,and...

文档格式:PDF | 页数:10 | 浏览次数:11 | 上传日期:2018-09-01 01:43:38 | 文档星级:
Mathematics of Deep LearningRené Vidal Joan Bruna Raja Giryes Stefano SoattoAbstract—Recently there has been a dramatic increase in theperformance of recognition systems due to the introduction ofdeep architectures for representation learning and classif i cation.However, the mathematical reasons for this success remainelusive. This tutorial will review recent work that aims toprovide a mathematical justif i cation for several properties ofdeep networks, such as global optimality, geometric stability,and invariance of the learned representations.I. I NTRODUCTIONDeep networks [1] are parametric models that perform se-quential operations on their input data. Each such operation,colloquially called a “layer”, consists of a linear transfor-mation, say, a convolution of its input, followed by a point-wise nonlinear “activation function”, e.g., a sigmoid. Deepnetworks have recently led to dramatic improvements inclassif ication performance in various applications in speechand natural language processing, and computer vision. Thecrucial property of deep networks that is believed to be theroot of their performance is that they have a large number oflayers as compared to classical neural networks; but there areother architectural modif ications such as rectif ied linear acti-vations (ReLUs) [2] and residual “shortcut” connections [3].Other major factors in their success is the availability ofmassive datasets, say, millions of images in datasets likeImageNet [4], and eff icient GPU computing hardware forsolving the resultant high-dimensional optimization problemwhich may have up to 100 million parameters.The empirical success of deep learning, especially con-volutional neural networks (CNNs) for image-based tasks,presents numerous puzzles to theoreticians. In particular,there are three key factors in deep learning, namely thearchitectures, regularization techniques and optimization al-gorithms, which are critical to train well-performing deepnetworks and understanding their necessity and interplay isessential if we are to unravel the secrets of their success.A. Approximation, depth, width and invariance propertiesAn important property in the design of a neural networkarchitecture is its ability to approximate arbitrary functionsof the input. But how does this ability depend on parametersof the architecture, such as its depth and width? Earlier workshows that neural networks with a single hidden layer andR. Vidal is with the Center for Imaging Science, Biomedical Engineering,Johns Hopkins University, Baltimore, USA rvidal@cis.jhu.eduJ. Bruna is with the Courant Institute of Mathematical Sciences, Centerfor Data Science, New York University, USA bruna@cims.nyu.eduR. Giryes is with the School of Electrical Engineering, Tel-Aviv Univer-sity, Tel Aviv, Israel raja@tauex.tau.ac.ilS. Soatto is with the Department of Computer Science, University ofCalifornia, Los Angeles, USA soatto@ucla.edusigmoidal activations are universal function approximators[5], [6], [7], [8]. However, the capacity of a wide and shallownetwork can be replicated by a deep network with signif icantimprovements in performance. One possible explanation isthat deeper architectures are able to better capture invariantproperties of the data compared to their shallow counterparts.In computer vision, for example, the category of an objectis invariant to changes in viewpoint, illumination, etc. Whilea mathematical analysis of why deep networks are able tocapture such invariances remains elusive, recent progresshas shed some light on this issue for certain sub-classesof deep networks. In particular, scattering networks [9] area class of deep networks whose convolutional f ilter banksare given by complex, multi-resolution wavelet families.As a result of this extra structure, they are provably stableand locally invariant signal representations, and reveal thefundamental role of geometry and stability that underpinsthe generalization performance of modern deep convolutionalnetwork architectures; see Section IV.B. Generalization and regularization propertiesAnother critical property of a neural network architectureis its ability to generalize from a small number of trainingexamples. Traditional results from statistical learning theory[10] show that the number of training examples needed toachieve good generalization grows polynomially with thesize of the network. In practice, however, deep networks aretrained with much fewer data than the number of parameters(N ? D regime) and yet they can be prevented from over-f itting using very simple (and seemingly counter-productive)regularization techniques like Dropout [11], which simplyfreezes a random subset of the parameters at each iteration.One possible explanation for this conundrum is that deeperarchitectures produce an embedding of the input data thatapproximately preserves the distance between data pointsin the same class, while increasing the separation betweenclasses. This tutorial will overview the recent work of [12],which uses tools from compressed sensing and dictionarylearning to prove that deep networks with random Gaussianweights perform a distance-preserving embedding of the datain which similar inputs are likely to have a similar output.These results provide insights into the metric learning prop-erties of the network and lead to bounds on the generalizationerror that are informed by the structure of the input data.C. Information-theoretic propertiesAnother key property of a network architecture is its abil-ity to produce a good “representation of the data”. Roughlyspeaking, a representation is any function of the input dataarXiv:1712.04741v1 [cs.LG] 13 Dec 2017

188宝金博页面版:关注我们

  • 新浪微博

关注188宝金博页面版公众号

188宝金博页面版
阅读
APP
阅读
返回
顶部
188宝金博页面版官网登录在线平台入口(2026已更新)—江苏协昌电子科技股份有限公司