Transfer language selection for zero-shot cross-lingual abusive language detection

Eronen, Juuso; Ptaszyński, Michał; Masui, Fumito; Arata, Masaki; Leliwa, Gniewosz; Wroczyński, Michał

doi:10.1016/j.ipm.2022.102981

Cited by 19 publications

(11 citation statements)

References 50 publications

(69 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…As a result, many studies employed Pre-trained multilingual word embeddings like FastText ( Bigoulaeva, Hangya & Fraser, 2021 ), MUSE ( Pamungkas & Patti, 2019 ; Deshpande, Farris & Kumar, 2022 ; Aluru et al, 2020 ; Bigoulaeva, Hangya & Fraser, 2021 ), or LASER ( Deshpande, Farris & Kumar, 2022 , Aluru et al, 2020 , Pelicon et al, 2021a ), and Vitiugin, Senarath & Purohit (2021) . Moreover, most of the research studies has focused on the use of pre-trained language models LLMs (basically as classifiers): BERT ( Vashistha & Zubiaga, 2021 , zahra El-Alami, Ouatik El Alaoui & En Nahnahi, 2022 ; Zia et al, 2022 ; Pamungkas, Basile & Patti, 2021a ), AraBERT (for Arabic data) ( zahra El-Alami, Ouatik El Alaoui & En Nahnahi, 2022 ), CseBERT (for English, Croatian and Slovenian data) ( Pelicon et al, 2021b ), as well as multilingual BERT models: ( Shi et al, 2022 ; Bhatia et al, 2021 ; Deshpande, Farris & Kumar, 2022 ; Aluru et al, 2020 ; zahra El-Alami, Ouatik El Alaoui & En Nahnahi, 2022 ; De la Peña Sarracén & Rosso, 2022 ; Tita & Zubiaga, 2021 ; Eronen et al, 2022 ; Ranasinghe & Zampieri, 2021a ; Ghadery & Moens, 2020 ; Pelicon et al, 2021b ; Awal et al, 2024 ; Montariol, Riabi & Seddah, 2022 ; Ahn et al, 2020a ; Bigoulaeva et al, 2022 , 2023 ; Pamungkas, Basile & Patti, 2021a ; Pelicon et al, 2021a ), DistilmBERT model ( Vitiugin, Senarath & Purohit, 2021 ), and RoBERTa ( Zia et al, 2022 ).…”

Section: Approaches On Multilingual Hate Speech Detectionmentioning

confidence: 99%

“…On the other hand, cross-lingual language models like XLM were also widely employed, where we found implementation of XLM-RoBERTa (XLM-R) ( Roy et al, 2021a ; Bhatia et al, 2021 ; Wang et al, 2020 ; De la Peña Sarracén & Rosso, 2022 ; Zia et al, 2022 ; Tita & Zubiaga, 2021 ; Ranasinghe & Zampieri, 2021b ; Dadu & Pant, 2020 ; Eronen et al, 2022 ; Ranasinghe & Zampieri, 2021a , 2020 ; Mozafari, Farahbakhsh & Crespi, 2022 ; Barbieri, Espinosa Anke & Camacho-Collados, 2022 ; Awal et al, 2024 ; Stappen, Brunn & Schuller, 2020 ), and both XLM-R and XLM-T ( Montariol, Riabi & Seddah, 2022 , Riabi, Montariol & Seddah, 2022 ). These approaches have all been shown to improve performance on tasks involving multilingual/cross-lingual hate speech detection because they are more likely able to capture semantic and syntactic features across languages thanks to their pre-training on multilingual large volumes of texts.…”

Section: Approaches On Multilingual Hate Speech Detectionmentioning

confidence: 99%

See 1 more Smart Citation

A survey on multi-lingual offensive language detection

Mnassri,

Farahbakhsh,

Chalehchaleh

et al. 2024

PeerJ Computer Science

View full text Add to dashboard Cite

The prevalence of offensive content on online communication and social media platforms is growing more and more common, which makes its detection difficult, especially in multilingual settings. The term “Offensive Language” encompasses a wide range of expressions, including various forms of hate speech and aggressive content. Therefore, exploring multilingual offensive content, that goes beyond a single language, focus and represents more linguistic diversities and cultural factors. By exploring multilingual offensive content, we can broaden our understanding and effectively combat the widespread global impact of offensive language. This survey examines the existing state of multilingual offensive language detection, including a comprehensive analysis on previous multilingual approaches, and existing datasets, as well as provides resources in the field. We also explore the related community challenges on this task, which include technical, cultural, and linguistic ones, as well as their limitations. Furthermore, in this survey we propose several potential future directions toward more efficient solutions for multilingual offensive language detection, enabling safer digital communication environment worldwide.

show abstract

Section: Approaches On Multilingual Hate Speech Detectionmentioning

confidence: 99%

Section: Approaches On Multilingual Hate Speech Detectionmentioning

confidence: 99%

A survey on multi-lingual offensive language detection

Mnassri,

Farahbakhsh,

Chalehchaleh

et al. 2024

PeerJ Computer Science

View full text Add to dashboard Cite

show abstract

“…Cyberbullying detection has gained considerable attention in recent years owing to the widespread use of social media platforms and online communication channels [7]. Researchers have explored various techniques and methodologies to identify and address cyberbullying among various languages and cultures [7,8,[31][32][33][34][35]. However, while numerous research efforts have introduced solutions to detect cyberbullying in high-resource languages such as English or Japanese, there is a limited number of studies that have extensively addressed cyberbullying detection in the low-resource languages, such as the Bangla language.…”

Section: Related Workmentioning

confidence: 99%

“…Through exposure to a variety of linguistic data, Multilingual BERT has the ability to process and convey the subtleties of various languages, including Bangla. Because of its strong architecture, it can handle a wide range of natural language processing tasks, such as the detection of cyberbullying [35], which makes it an invaluable tool for multilingual text analysis.…”

Section: Multilingual Bertmentioning

confidence: 99%

Exhaustive Study into Machine Learning and Deep Learning Methods for Multilingual Cyberbullying Detection in Bangla and Chittagonian Texts

Mahmud,

Ptaszynski,

Masui

2024

Electronics

Self Cite

View full text Add to dashboard Cite

Cyberbullying is a serious problem in online communication. It is important to find effective ways to detect cyberbullying content to make online environments safer. In this paper, we investigated the identification of cyberbullying contents from the Bangla and Chittagonian languages, which are both low-resource languages, with the latter being an extremely low-resource language. In the study, we used both traditional baseline machine learning methods, as well as a wide suite of deep learning methods especially focusing on hybrid networks and transformer-based multilingual models. For the data, we collected over 5000 both Bangla and Chittagonian text samples from social media. Krippendorff’s alpha and Cohen’s kappa were used to measure the reliability of the dataset annotations. Traditional machine learning methods used in this research achieved accuracies ranging from 0.63 to 0.711, with SVM emerging as the top performer. Furthermore, employing ensemble models such as Bagging with 0.70 accuracy, Boosting with 0.69 accuracy, and Voting with 0.72 accuracy yielded promising results. In contrast, deep learning models, notably CNN, achieved accuracies ranging from 0.69 to 0.811, thus outperforming traditional ML approaches, with CNN exhibiting the highest accuracy. We also proposed a series of hybrid network-based models, including BiLSTM+GRU with an accuracy of 0.799, CNN+LSTM with 0.801 accuracy, CNN+BiLSTM with 0.78 accuracy, and CNN+GRU with 0.804 accuracy. Notably, the most complex model, (CNN+LSTM)+BiLSTM, attained an accuracy of 0.82, thus showcasing the efficacy of hybrid architectures. Furthermore, we explored transformer-based models, such as XLM-Roberta with 0.841 accuracy, Bangla BERT with 0.822 accuracy, Multilingual BERT with 0.821 accuracy, BERT with 0.82 accuracy, and Bangla ELECTRA with 0.785 accuracy, which showed significantly enhanced accuracy levels. Our analysis demonstrates that deep learning methods can be highly effective in addressing the pervasive issue of cyberbullying in several different linguistic contexts. We show that transformer models can efficiently circumvent the language dependence problem that plagues conventional transfer learning methods. Our findings suggest that hybrid approaches and transformer-based embeddings can effectively tackle the problem of cyberbullying across online platforms.

show abstract

“…Cross-Lingual Abusive Language Detection. In recent years, crosslingual abusive language detection has gained increasing attention in zeroshot (Eronen et al, 2022) and few-shot (Mozafari et al, 2022) transfer.…”

Section: Background and Related Workmentioning

confidence: 99%

On the Keyword Extraction and Bias Analysis, Graph-based Exploration and Data Augmentation for Abusive Language Detection in Low-Resource Settings

Peña Sarracén

View full text Add to dashboard Cite

Many thanks go to my colleagues and fellow researchers at the Universitat Politècnica de València, whose collaboration and support have been both enjoyable and intellectually enriching. The synergy between academia and industry fostered by Symanto has been instrumental in achieving the objectives of this thesis.I extend my appreciation to the reviewers of this thesis Dr. Arkaitz Zubiaga (Queen Mary University of London), Prof. Rafael Valencia (Universidad de Murcia), and Dr. Tommaso Caselli (Rijksuniveristeit Groningen); thank you very much for your work on my thesis. I would also like to thank the members of the PhD committee of this thesis Dr. Arkaitz Zubiaga (Queen Mary University of London), Dra. Sara Tonelli (Fondazione Bruno Kessler), and Dra. Aitziber Atutxa (University of the Basque Country).Special thanks go to my family for their tireless encouragement and understanding. Their love and support have been my pillars throughout this challenging yet rewarding endeavor. Also to the friends who have taken an interest in my emotional and physical well-being during my doctoral research.Last but not least, I would like to thank all the participants and individuals who generously contributed their time. I am truly grateful.

show abstract

Transfer language selection for zero-shot cross-lingual abusive language detection

Cited by 19 publications

References 50 publications

A survey on multi-lingual offensive language detection

A survey on multi-lingual offensive language detection

Exhaustive Study into Machine Learning and Deep Learning Methods for Multilingual Cyberbullying Detection in Bangla and Chittagonian Texts

On the Keyword Extraction and Bias Analysis, Graph-based Exploration and Data Augmentation for Abusive Language Detection in Low-Resource Settings

Contact Info

Product

Resources

About