Hierarchical Deep Learning for Arabic Dialect Identification

Francony, Gael de; Guichard, Victor; Joshi, Praveen; Afli, Haithem; Bouchekif, Abdessalam

doi:10.18653/v1/w19-4631

Cited by 5 publications

(2 citation statements)

References 6 publications

(6 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Many attempts have been proposed in the area of automatic dialect identification (ADI), and early uses are based on dictionaries, rules, and language modeling [5][6][7][8][9][10]; more recently, a shift was made toward employing machine learning techniques [11][12][13][14][15][16][17][18][19][20][21][22][23][24], deep learning approaches [25][26][27][28][29][30][31][32][33][34][35][36][37][38][39][40], and transfer learning methods [41][42][43][44][45][46][47][48][49]. Many of these investigations utilize prominent and accessible datasets, such as MADAR [49], NADI [50][51]…”

Section: Introductionmentioning

confidence: 99%

Enhancing Arabic Dialect Detection on Social Media: A Hybrid Model with an Attention Mechanism

Yafooz

2024

Information

View full text Add to dashboard Cite

Recently, the widespread use of social media and easy access to the Internet have brought about a significant transformation in the type of textual data available on the Web. This change is particularly evident in Arabic language usage, as the growing number of users from diverse domains has led to a considerable influx of Arabic text in various dialects, each characterized by differences in morphology, syntax, vocabulary, and pronunciation. Consequently, researchers in language recognition and natural language processing have become increasingly interested in identifying Arabic dialects. Numerous methods have been proposed to recognize this informal data, owing to its crucial implications for several applications, such as sentiment analysis, topic modeling, text summarization, and machine translation. However, Arabic dialect identification is a significant challenge due to the vast diversity of the Arabic language in its dialects. This study introduces a novel hybrid machine and deep learning model, incorporating an attention mechanism for detecting and classifying Arabic dialects. Several experiments were conducted using a novel dataset that collected information from user-generated comments from Twitter of Arabic dialects, namely, Egyptian, Gulf, Jordanian, and Yemeni, to evaluate the effectiveness of the proposed model. The dataset comprises 34,905 rows extracted from Twitter, representing an unbalanced data distribution. The data annotation was performed by native speakers proficient in each dialect. The results demonstrate that the proposed model outperforms the performance of long short-term memory, bidirectional long short-term memory, and logistic regression models in dialect classification using different word representations as follows: term frequency-inverse document frequency, Word2Vec, and global vector for word representation.

show abstract

Section: Introductionmentioning

confidence: 99%

Enhancing Arabic Dialect Detection on Social Media: A Hybrid Model with an Attention Mechanism

Yafooz

2024

Information

View full text Add to dashboard Cite

show abstract

“…MADAR has been established as an important corpus for the task, serving as a benchmark for multi-task learning (Seelawi et al, 2021), as well as a Shared Task corpus (Bouamor et al, 2019), and as a subject of independent research (Baimukan et al, 2022). Despite several attempts to develop models using deep neural networks (Lippincott et al, 2019;de Francony et al, 2019) and pre-trained Transformer-based language models (Inoue et al, 2021), the current state-of-theart approach remains a statistical machine learning model with surface-level feature representation, specifically the Multinomial Naive Bayes (MNB) model introduced by .…”

Section: Introductionmentioning

confidence: 99%

Arabic dialect identification: An in-depth error analysis on the MADAR parallel corpus

Olsen,

Touileb,

Velldal

2023

Proceedings of ArabicNLP 2023

View full text Add to dashboard Cite

This paper provides a systematic analysis and comparison of the performance of state-of-theart models on the task of fine-grained Arabic dialect identification using the MADAR parallel corpus. We test approaches based on pretrained transformer language models in addition to Naive Bayes models with a rich set of various features. Through a comprehensive data-and error analysis, we provide valuable insights into the strengths and weaknesses of both approaches. We discuss which dialects are more challenging to differentiate, and identify potential sources of errors. Our analysis reveals an important problem with identical sentences across dialect classes in the test set of the MADAR-26 corpus, which may confuse any classifier. We also show that none of the tested approaches captures the subtle distinctions between closely related dialects.

show abstract