Gaussian Constrained Attention Network for Scene Text Recognition

by   Zhi Qiao, et al.

Scene text recognition has been a hot topic in computer vision. Recent methods adopt the attention mechanism for sequence prediction which achieve convincing results. However, we argue that the existing attention mechanism faces the problem of attention diffusion, in which the model may not focus on a certain character area. In this paper, we propose Gaussian Constrained Attention Network to deal with this problem. It is a 2D attention-based method integrated with a novel Gaussian Constrained Refinement Module, which predicts an additional Gaussian mask to refine the attention weights. Different from adopting an additional supervision on the attention weights simply, our proposed method introduces an explicit refinement. In this way, the attention weights will be more concentrated and the attention-based recognition network achieves better performance. The proposed Gaussian Constrained Refinement Module is flexible and can be applied to existing attention-based methods directly. The experiments on several benchmark datasets demonstrate the effectiveness of our proposed method. Our code has been available at



There are no comments yet.


page 1

page 6


Focusing Attention: Towards Accurate Text Recognition in Natural Images

Scene text recognition has been a hot research topic in computer vision ...

A Multi-Object Rectified Attention Network for Scene Text Recognition

Irregular text is widely used. However, it is considerably difficult to ...

A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolution

Scene text image super-resolution aims to increase the resolution and re...

Implicit Rate-Constrained Optimization of Non-decomposable Objectives

We consider a popular family of constrained optimization problems arisin...

Rethinking Text Segmentation: A Novel Dataset and A Text-Specific Refinement Approach

Text segmentation is a prerequisite in many real-world text-related task...

Generic Event Boundary Detection Challenge at CVPR 2021 Technical Report: Cascaded Temporal Attention Network (CASTANET)

This report presents the approach used in the submission of Generic Even...

Causal Attention for Vision-Language Tasks

We present a novel attention mechanism: Causal Attention (CATT), to remo...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.