Skip to main navigation Skip to search Skip to main content

Activation function design for deep networks: linearity and effective initialisation

  • University of Oxford
  • The Alan Turing Institute

Research output: Contribution to journalArticlepeer-review

20   Link opens in a new tab Citations (SciVal)

Abstract

The activation function deployed in a deep neural network has great influence on the performance of the network at initialisation, which in turn has implications for training. In this paper we study how to avoid two problems at initialisation identified in prior works: rapid convergence of pairwise input correlations, and vanishing and exploding gradients. We prove that both these problems can be avoided by choosing an activation function possessing a sufficiently large linear region around the origin, relative to the bias variance σb2 of the network's random initialisation. We demonstrate empirically that using such activation functions leads to tangible benefits in practice, both in terms of test and training accuracy and in terms of training time. Furthermore, we observe that the shape of the nonlinear activation outside the linear region appears to have a relatively limited impact on training. Finally, our results also allow us to train networks in a new hyperparameter regime, with a much larger bias variance than has previously been possible.

Original languageEnglish
Pages (from-to)117-154
Number of pages38
JournalApplied and Computational Harmonic Analysis
Volume59
Early online date25 Apr 2022
DOIs
Publication statusPublished - 31 Jul 2022

Bibliographical note

Publisher Copyright:
© 2022 Elsevier Inc.

Keywords

  • Activation function design
  • Deep learning
  • Initialisation
  • Random networks

ASJC Scopus subject areas

  • Applied Mathematics

Fingerprint

Dive into the research topics of 'Activation function design for deep networks: linearity and effective initialisation'. Together they form a unique fingerprint.

Cite this