From Nodes to Networks: Evolving Recurrent Neural Networks

03/12/2018
by   Aditya Rawal, et al.
0

Gated recurrent networks such as those composed of Long Short-Term Memory (LSTM) nodes have recently been used to improve state of the art in many sequential processing tasks such as speech recognition and machine translation. However, the basic structure of the LSTM node is essentially the same as when it was first conceived 25 years ago. Recently, evolutionary and reinforcement learning mechanisms have been employed to create new variations of this structure. This paper proposes a new method, evolution of a tree-based encoding of the gated memory nodes, and shows that it makes it possible to explore new variations more effectively than other methods. The method discovers nodes with multiple recurrent paths and multiple memory cells, which lead to significant improvement in the standard language modeling benchmark task. The paper also shows how the search process can be speeded up by training an LSTM network to estimate performance of candidate structures, and by encouraging exploration of novel solutions. Thus, evolutionary design of complex neural network structures promises to improve performance of deep learning architectures beyond human ability to do so.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/28/2016

Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition

Recurrent neural networks (RNNs) have shown clear superiority in sequenc...
research
03/08/2017

Interpretable Structure-Evolving LSTM

This paper develops a general framework for learning interpretable data ...
research
11/07/2017

Cortical microcircuits as gated-recurrent neural networks

Cortical circuits exhibit intricate recurrent architectures that are rem...
research
09/11/2020

Applications of Deep Neural Networks

Deep learning is a group of exciting new technologies for neural network...
research
10/09/2019

Kernel-Based Approaches for Sequence Modeling: Connections to Neural Methods

We investigate time-dependent data analysis from the perspective of recu...
research
12/20/2017

A Flexible Approach to Automated RNN Architecture Generation

The process of designing neural architectures requires expert knowledge ...
research
03/13/2015

LSTM: A Search Space Odyssey

Several variants of the Long Short-Term Memory (LSTM) architecture for r...

Please sign up or login with your details

Forgot password? Click here to reset