Dia AbuZeina scite author profile

One of the problems in the speech recognition of Modern Standard Arabic (MSA) is the cross-word pronunciation variation. Cross-word pronunciation variations alter the phonetic spelling of words beyond their listed forms in the phonetic dictionary, leading to a number of Out-Of-Vocabulary (OOV) wordforms. This paper presents a knowledge-based approach to model cross-word pronunciation variation at both phonetic dictionary and language model levels. The proposed approach is based on modeling cross-word pronunciation variation by expanding the phonetic dictionary and corpus transcription. The Baseline system contains a phonetic dictionary of 14,234 words from a 5.4 hours corpus of Arabic broadcast news. The expanded dictionary contains 15,873 words. Also, the corpus transcription is expanded according to the applied Arabic phonological rules. Using Carnegie Mellon University (CMU) Sphinx speech recognition engine, the Enhanced system achieved Word Error Rate (WER) of 9.91% on a test set of fully discretized transcription of about 1.1 hours of Arabic broadcast news. The WER is enhanced by 2.3% compared to the Baseline system.

show abstract

Stemming impact on Arabic text categorization performance: A survey

Al-Anzi

AbuZeina

2015

View full text Add to dashboard Cite

Synopsis on Arabic speech recognition

Al-Anzi

AbuZeina

2022

Ain Shams Engineering Journal

View full text Add to dashboard Cite

Arabic Speech Recognition Systems

AbuZeina

Elshafei

2011

View full text Add to dashboard Cite

Literature Survey of Arabic Speech Recognition

Al-Anzi

AbuZeina

2018

View full text Add to dashboard Cite

12 3 4

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Dia AbuZeina

Toward an enhanced Arabic text classification using cosine similarity and Latent Semantic Indexing

Beyond vector space model for hierarchical Arabic text classification: A Markov chain approach

Employing fisher discriminant analysis for Arabic text classification

Cross-word Arabic pronunciation variation modeling for speech recognition

Stemming impact on Arabic text categorization performance: A survey

Synopsis on Arabic speech recognition

Arabic Speech Recognition Systems

Literature Survey of Arabic Speech Recognition

Contact Info

Product

Resources

About