New Results for the Text Recognition of Arabic Maghribī Manuscripts – Managing an Under-resourced Script

11/29/2022
by   Lucas Noëmie, et al.
0

HTR models development has become a conventional step for digital humanities projects. The performance of these models, often quite high, relies on manual transcription and numerous handwritten documents. Although the method has proven successful for Latin scripts, a similar amount of data is not yet achievable for scripts considered poorly-endowed, like Arabic scripts. In that respect, we are introducing and assessing a new modus operandi for HTR models development and fine-tuning dedicated to the Arabic Maghribī scripts. The comparison between several state-of-the-art HTR demonstrates the relevance of a word-based neural approach specialized for Arabic, capable to achieve an error rate below 5 perspectives for Arabic scripts processing and more generally for poorly-endowed languages processing. This research is part of the development of RASAM dataset in partnership with the GIS MOMM and the BULAC.

READ FULL TEXT

page 8

page 15

page 16

page 20

page 21

page 23

page 28

page 29

research
09/18/2020

An Efficient Language-Independent Multi-Font OCR for Arabic Script

Optical Character Recognition (OCR) is the process of extracting digitiz...
research
10/15/2018

Diacritization of Maghrebi Arabic Sub-Dialects

Diacritization process attempt to restore the short vowels in Arabic wri...
research
06/29/2021

New Arabic Medical Dataset for Diseases Classification

The Arabic language suffers from a great shortage of datasets suitable f...
research
02/04/2020

Arabic Diacritic Recovery Using a Feature-Rich biLSTM Model

Diacritics (short vowels) are typically omitted when writing Arabic text...
research
10/19/2022

Arabic Word-level Readability Visualization for Assisted Text Simplification

This demo paper presents a Google Docs add-on for automatic Arabic word-...
research
07/30/2019

EdgeNet: A novel approach for Arabic numeral classification

Despite the importance of handwritten numeral classification, a robust a...
research
09/29/2021

Improving Arabic Diacritization by Learning to Diacritize and Translate

We propose a novel multitask learning method for diacritization which tr...

Please sign up or login with your details

Forgot password? Click here to reset