Adaptive-saturated RNN: Remember more with less instability

04/24/2023

∙

Orthogonal parameterization is a compelling solution to the vanishing gradient problem (VGP) in recurrent neural networks (RNNs). With orthogonal parameters and non-saturated activation functions, gradients in such models are constrained to unit norms. On the other hand, although the traditional vanilla RNNs are seen to have higher memory capacity, they suffer from the VGP and perform badly in many applications. This work proposes Adaptive-Saturated RNNs (asRNN), a variant that dynamically adjusts its saturation level between the two mentioned approaches. Consequently, asRNN enjoys both the capacity of a vanilla RNN and the training stability of orthogonal RNNs. Our experiments show encouraging results of asRNN on challenging sequence learning benchmarks compared to several strong competitors. The research code is accessible at https://github.com/ndminhkhoi46/asRNN/.

READ FULL TEXT

Adaptive-saturated RNN: Remember more with less instability

MomentumRNN: Integrating Momentum into Recurrent Neural Networks

Deep Independently Recurrent Neural Network (IndRNN)

Dilated Recurrent Neural Networks

Improved memory in recurrent neural networks with sequential non-normal dynamics

Capacity and Trainability in Recurrent Neural Networks

An Adaptive Stochastic Nesterov Accelerated Quasi Newton Method for Training RNNs

Non-normal Recurrent Neural Network (nnRNN): learning long time dependencies while improving expressivity with transient dynamics

Adaptive-saturated RNN: Remember more with less instability

Related Research

MomentumRNN: Integrating Momentum into Recurrent Neural Networks

Deep Independently Recurrent Neural Network (IndRNN)

Dilated Recurrent Neural Networks

Improved memory in recurrent neural networks with sequential non-normal dynamics

Capacity and Trainability in Recurrent Neural Networks

An Adaptive Stochastic Nesterov Accelerated Quasi Newton Method for Training RNNs

Non-normal Recurrent Neural Network (nnRNN): learning long time dependencies while improving expressivity with transient dynamics