Improving Multi-Head Attention with Capsule Networks

08/31/2019
by   Shuhao Gu, et al.
0

Multi-head attention advances neural machine translation by working out multiple versions of attention in different subspaces, but the neglect of semantic overlapping between subspaces increases the difficulty of translation and consequently hinders the further improvement of translation performance. In this paper, we employ capsule networks to comb the information from the multiple heads of the attention so that similar information can be clustered and unique information can be reserved. To this end, we adopt two routing mechanisms of Dynamic Routing and EM Routing, to fulfill the clustering and separating. We conducted experiments on Chinese-to-English and English-to-German translation tasks and got consistent improvements over the strong Transformer baseline.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/30/2020

Capsule-Transformer for Neural Machine Translation

Transformer hugely benefits from its key design of the multi-head self-a...
research
10/24/2018

Multi-Head Attention with Disagreement Regularization

Multi-head attention is appealing for the ability to jointly attend to i...
research
02/15/2019

Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement

With the promising progress of deep neural networks, layer aggregation h...
research
04/05/2019

Information Aggregation for Multi-Head Attention with Routing-by-Agreement

Multi-head attention is appealing for its ability to jointly extract dif...
research
09/11/2018

On The Alignment Problem In Multi-Head Attention-Based Neural Machine Translation

This work investigates the alignment problem in state-of-the-art multi-h...
research
11/01/2018

Towards Linear Time Neural Machine Translation with Capsule Networks

In this study, we first investigate a novel capsule network with dynamic...
research
04/21/2019

Dynamic Past and Future for Neural Machine Translation

Previous studies have shown that neural machine translation (NMT) models...

Please sign up or login with your details

Forgot password? Click here to reset