A Quranic Dataset for Text Recognition

Al-Sheikh, Idris; Mohd, Masnizah

doi:10.4108/eai.18-7-2019.2287842

Cited by 2 publications

(1 citation statement)

References 6 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…The text recognition system must recognise the whole word, but the Arabic language is a cursive language where the characters are connected to construct a word or subword. Thus, [26,27] introduced the subword dataset where they built the Arabic word from more than one group, such as the word Quran ‫قران(‬ (.…”

Section: Introductionmentioning

confidence: 99%

Quranic Optical Text Recognition Using Deep Learning Models

et al. 2021

Self Cite

View full text Add to dashboard Cite

A Quranic optical character recognition (OCR) system based on convolutional neural network (CNN) followed by recurrent neural network (RNN) is introduced in this work. Six deep learning models are built to study the effect of different representations of the input and output, and the accuracy and performance of the models, and compare long short-term memory (LSTM) and gated recurrent unit (GRU). A new Quranic OCR dataset is developed based on the most famous printed version of the Holy Quran (Mushaf Al-Madinah), and a page and line-text image with the corresponding labels is prepared. This work's contribution is a Quranic OCR model capable of recognising the Quranic image's diacritic text. A better performance in word recognition rate (WRR) and character recognition rate (CRR) is achieved in the experiments. The LSTM and GRU are compared in the Arabic text recognition domain. In addition, a public database is built for research purposes in Arabic text recognition that contains the diacritics and the Uthmanic script, and is large enough to be used with the deep learning models. The outcome of this work shows that the proposed system obtains an accuracy of 98% on the validation data, and a WRR of 95% and a CRR of 99% in the test dataset.

show abstract

Section: Introductionmentioning

confidence: 99%