Chinese Text in the Wild

02/28/2018
by   Tai-Ling Yuan, et al.
0

We introduce Chinese Text in the Wild, a very large dataset of Chinese text in street view images. While optical character recognition (OCR) in document images is well studied and many commercial tools are available, detection and recognition of text in natural images is still a challenging problem, especially for more complicated character sets such as Chinese text. Lack of training data has always been a problem, especially for deep learning methods which require massive training data. In this paper we provide details of a newly created dataset of Chinese text with about 1 million Chinese characters annotated by experts in over 30 thousand street view images. This is a challenging dataset with good diversity. It contains planar text, raised text, text in cities, text in rural areas, text under poor illumination, distant text, partially occluded text, etc. For each character in the dataset, the annotation includes its underlying character, its bounding box, and 6 attributes. The attributes indicate whether it has complex background, whether it is raised, whether it is handwritten or printed, etc. The large size and diversity of this dataset make it suitable for training robust neural networks for various tasks, particularly detection and recognition. We give baseline results using several state-of-the-art networks, including AlexNet, OverFeat, Google Inception and ResNet for character recognition, and YOLOv2 for character detection in images. Overall Google Inception has the best performance on recognition with 80.5 while YOLOv2 achieves an mAP of 71.0 trained models will all be publicly available on the website.

READ FULL TEXT

page 4

page 5

page 6

page 7

page 9

research
03/25/2019

ShopSign: a Diverse Scene Text Dataset of Chinese Shop Signs in Street Views

In this paper, we introduce the ShopSign dataset, which is a newly devel...
research
04/07/2016

A CNN Based Scene Chinese Text Recognition Algorithm With Synthetic Data Engine

Scene text recognition plays an important role in many computer vision a...
research
01/26/2016

COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

This paper describes the COCO-Text dataset. In recent years large-scale ...
research
09/17/2019

Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning

Most existing text reading benchmarks make it difficult to evaluate the ...
research
12/27/2022

A Comprehensive Gold Standard and Benchmark for Comics Text Detection and Recognition

This study focuses on improving the optical character recognition (OCR) ...
research
09/21/2020

PP-OCR: A Practical Ultra Lightweight OCR System

The Optical Character Recognition (OCR) systems have been widely used in...
research
11/23/2022

Indian Commercial Truck License Plate Detection and Recognition for Weighbridge Automation

Detection and recognition of a licence plate is important when automatin...

Please sign up or login with your details

Forgot password? Click here to reset