Morphological Reconstruction for Word Level Script Identification

06/25/2011
by   B. V. Dhandra, et al.
0

A line of a bilingual document page may contain text words in regional language and numerals in English. For Optical Character Recognition (OCR) of such a document page, it is necessary to identify different script forms before running an individual OCR system. In this paper, we have identified a tool of morphological opening by reconstruction of an image in different directions and regional descriptors for script identification at word level, based on the observation that every text has a distinct visual appearance. The proposed system is developed for three Indian major bilingual documents, Kannada, Telugu and Devnagari containing English numerals. The nearest neighbour and k-nearest neighbour algorithms are applied to classify new word images. The proposed algorithm is tested on 2625 words with various font styles and sizes. The results obtained are quite encouraging

READ FULL TEXT
research
05/10/2012

Discrimination of English to other Indian languages (Kannada and Hindi) for OCR system

India is a multilingual multi-script country. In every state of India th...
research
10/22/2013

Word Spotting in Cursive Handwritten Documents using Modified Character Shape Codes

There is a large collection of Handwritten English paper documents of Hi...
research
12/11/2022

Extending TrOCR for Text Localization-Free OCR of Full-Page Scanned Receipt Images

Digitization of scanned receipts aims to extract text from receipt image...
research
02/13/2022

Omnifont Persian OCR System Using Primitives

In this paper, we introduce a model-based omnifont Persian OCR system. T...
research
07/29/2016

Labeling of Query Words using Conditional Random Field

This paper describes our approach on Query Word Labeling as an attempt i...
research
05/31/2021

Pho(SC)Net: An Approach Towards Zero-shot Word Image Recognition in Historical Documents

Annotating words in a historical document image archive for word image r...
research
01/27/2017

Document Decomposition of Bangla Printed Text

Today all kind of information is getting digitized and along with all th...

Please sign up or login with your details

Forgot password? Click here to reset