Video Text Tracking With a Spatio-Temporal Complementary Model

11/09/2021
by   Yuzhe Gao, et al.
0

Text tracking is to track multiple texts in a video,and construct a trajectory for each text. Existing methodstackle this task by utilizing the tracking-by-detection frame-work, i.e., detecting the text instances in each frame andassociating the corresponding text instances in consecutiveframes. We argue that the tracking accuracy of this paradigmis severely limited in more complex scenarios, e.g., owing tomotion blur, etc., the missed detection of text instances causesthe break of the text trajectory. In addition, different textinstances with similar appearance are easily confused, leadingto the incorrect association of the text instances. To this end,a novel spatio-temporal complementary text tracking model isproposed in this paper. We leverage a Siamese ComplementaryModule to fully exploit the continuity characteristic of the textinstances in the temporal dimension, which effectively alleviatesthe missed detection of the text instances, and hence ensuresthe completeness of each text trajectory. We further integratethe semantic cues and the visual cues of the text instance intoa unified representation via a text similarity learning network,which supplies a high discriminative power in the presence oftext instances with similar appearance, and thus avoids the mis-association between them. Our method achieves state-of-the-art performance on several public benchmarks. The source codeis available at https://github.com/lsabrinax/VideoTextSCM.

READ FULL TEXT

page 1

page 3

page 4

page 7

page 8

page 12

research
05/07/2021

MOTR: End-to-End Multiple-Object Tracking with TRansformer

The key challenge in multiple-object tracking (MOT) task is temporal mod...
research
11/14/2022

Discovering A Variety of Objects in Spatio-Temporal Human-Object Interactions

Spatio-temporal Human-Object Interaction (ST-HOI) detection aims at dete...
research
08/20/2019

An End-to-end Video Text Detector with Online Tracking

Video text detection is considered as one of the most difficult tasks in...
research
03/20/2022

End-to-End Video Text Spotting with Transformer

Recent video text spotting methods usually require the three-staged pipe...
research
11/09/2018

Multiple People Tracking Using Hierarchical Deep Tracklet Re-identification

The task of multiple people tracking in monocular videos is challenging ...
research
07/29/2020

MessyTable: Instance Association in Multiple Camera Views

We present an interesting and challenging dataset that features a large ...
research
11/15/2021

Tracking People with 3D Representations

We present a novel approach for tracking multiple people in video. Unlik...

Please sign up or login with your details

Forgot password? Click here to reset