Towards Spatio-Temporal Video Scene Text Detection via Temporal Clustering

11/19/2020
by   Yuanqiang Cai, et al.
0

With only bounding-box annotations in the spatial domain, existing video scene text detection (VSTD) benchmarks lack temporal relation of text instances among video frames, which hinders the development of video text-related applications. In this paper, we systematically introduce a new large-scale benchmark, named as STVText4, a well-designed spatial-temporal detection metric (STDM), and a novel clustering-based baseline method, referred to as Temporal Clustering (TC). STVText4 opens a challenging yet promising direction of VSTD, termed as ST-VSTD, which targets at simultaneously detecting video scene texts in both spatial and temporal domains. STVText4 contains more than 1.4 million text instances from 161,347 video frames of 106 videos, where each instance is annotated with not only spatial bounding box and temporal range but also four intrinsic attributes, including legibility, density, scale, and lifecycle, to facilitate the community. With continuous propagation of identical texts in the video sequence, TC can accurately output the spatial quadrilateral and temporal range of the texts, which sets a strong baseline for ST-VSTD. Experiments demonstrate the efficacy of our method and the great academic and practical value of the STVText4. The dataset and code will be available soon.

READ FULL TEXT

page 1

page 3

page 5

page 7

research
04/01/2020

Spatio-Temporal Action Detection with Multi-Object Interaction

Spatio-temporal action detection in videos requires localizing the actio...
research
01/06/2021

Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in Videos

Segmenting objects in videos is a fundamental computer vision task. The ...
research
05/25/2021

ST-HOI: A Spatial-Temporal Baseline for Human-Object Interaction Detection in Videos

Detecting human-object interactions (HOI) is an important step toward a ...
research
07/04/2022

Fast Vehicle Detection and Tracking on Fisheye Traffic Monitoring Video using CNN and Bounding Box Propagation

We design a fast car detection and tracking algorithm for traffic monito...
research
03/08/2019

Efficient Video Scene Text Spotting: Unifying Detection, Tracking, and Recognition

This paper proposes an unified framework for efficiently spotting scene ...
research
09/16/2022

An Attention-guided Multistream Feature Fusion Network for Localization of Risky Objects in Driving Videos

Detecting dangerous traffic agents in videos captured by vehicle-mounted...
research
09/17/2022

Spatial-Temporal Deep Embedding for Vehicle Trajectory Reconstruction from High-Angle Video

Spatial-temporal Map (STMap)-based methods have shown great potential to...

Please sign up or login with your details

Forgot password? Click here to reset