Proceedings of the 11th Forum for Information Retrieval Evaluation 2019
|
Sign up to set email alerts
Language Identification of Bengali-English Code-Mixed Data using Character & Phonetic based LSTM Models
Search citation statements
Order By: Relevance
Paper Sections
Select...
5
1
0
0
Citation Types
0
7
0
0
Year Published
Range
2021
20212024
2024Publication Types
Select...
5
Relationship
0
5
Authors
Journals
Cited by 5 publications
(7 citation statements)
References 6 publications
0
7
0
0
Order By: Relevance
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Singh et al [20] build an automatic named entity recognition (NER) system for Hindi-English code-mixed data and the proposed system outperformed with 33.18% of F1-score in comparison with existing baseline systems. However, this work can be extended to build natural language processing (NLP) [21] presented a supervised learning model for word-level language identification in Bengali-English code-mixed data. Two types of word encoding methods such as character and phonetic are used along with stacking and threshold techniques.…”
Section: Issn: 2088-8708
mentioning
confidence: 99%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Singh et al [20] build an automatic named entity recognition (NER) system for Hindi-English code-mixed data and the proposed system outperformed with 33.18% of F1-score in comparison with existing baseline systems. However, this work can be extended to build natural language processing (NLP) [21] presented a supervised learning model for word-level language identification in Bengali-English code-mixed data. Two types of word encoding methods such as character and phonetic are used along with stacking and threshold techniques.…”
Section: Issn: 2088-8708
mentioning
confidence: 99%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Another RNN variation technique, LSTM, has shown satisfactory performance in identifying Hindi-English and Bengali-English code-mixed text [27,52,54]. In [52], the LSTM architecture could give a high average F1 score of 93.4% and an average accuracy of 96.1% across the three classes.…”
Section: 1) Machine Learning Approach
mentioning
confidence: 99%
“…The following we identified some non-standard words encountered from the investigated papers. We categorised the non-standard words into four types, such as non-standard spelling [7,15,56], abbreviated words [3,37,39,45,49,56,64], exaggerated words [3,7,27,39,45,47,[49][50][51]64], and mixing characters with numbers or special characters [3,27,39,50]. Table 6 describes some examples of non-standard words found in code-mixed text LID.…”
Section: ) Non-standard Words
mentioning
confidence: 99%
“…Non-standard spelling [7,56] Prends or prenzz (friends), plis (please), kalo for 'kalau' (Indonesian language, meaning 'if' in English) Mixing word and numeric or special characters [3,7,27,39,50] ri8 (right), 2morrow (tomorrow), ni8t (night), orang2 (Indonesian language, meaning people in English) Word exaggeration [3,7,27,39,45,47,[49][50][51]64] goood (good), Pleasssseee (please), cooool (cool), helloooo (hello) Abbreviated words [3,39,45,49,56,64] bght (brought or bought), tkt (ticket), flm (film), TC (take care)…”
Section: Type Of Non-standard Word Example
mentioning
confidence: 99%
“…Austronesian & Germanic Indonesian-English [7] Malay-English [57] Dravidian & Germanic Kannada-English [44] Malayalam-English [47,64] Tamil-English [47,51] Telugu-English [50] Germanic & Trans Eurasian Dutch-Turkish [46] Germanic & Germanic Dutch-English [20,46,49] Dutch-Limburgish [43] German-English [46] Indo-Aryan & Germanic Assamese-English [21] Assamese-Hindi-Bengali-English [18,63] Bengali-English [11,27,54] Bengali-Hindi-English [54] Gujarati-English [59] Gujarati-Hindi-English [58] Hindi-English [11,15,37,48,52,54,60,61,65] Konkani-English [45] Punjabi-English [16] Sinhala-English [3,56] Italic & Germanic French-English [46] French-Italian-Spanish-English [55] Portuguese-English [46] Spanish-English [15,38,41,46,49] Italic & American Spanish-Wixarika…”
Section: Language Family Combination
mentioning
confidence: 99%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Singh et al [20] build an automatic named entity recognition (NER) system for Hindi-English code-mixed data and the proposed system outperformed with 33.18% of F1-score in comparison with existing baseline systems. However, this work can be extended to build natural language processing (NLP) [21] presented a supervised learning model for word-level language identification in Bengali-English code-mixed data. Two types of word encoding methods such as character and phonetic are used along with stacking and threshold techniques.…”
Section: Issn: 2088-8708
mentioning
confidence: 99%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Another RNN variation technique, LSTM, has shown satisfactory performance in identifying Hindi-English and Bengali-English code-mixed text [27,52,54]. In [52], the LSTM architecture could give a high average F1 score of 93.4% and an average accuracy of 96.1% across the three classes.…”
Section: 1) Machine Learning Approach
mentioning
confidence: 99%
“…The following we identified some non-standard words encountered from the investigated papers. We categorised the non-standard words into four types, such as non-standard spelling [7,15,56], abbreviated words [3,37,39,45,49,56,64], exaggerated words [3,7,27,39,45,47,[49][50][51]64], and mixing characters with numbers or special characters [3,27,39,50]. Table 6 describes some examples of non-standard words found in code-mixed text LID.…”
Section: ) Non-standard Words
mentioning
confidence: 99%
“…Non-standard spelling [7,56] Prends or prenzz (friends), plis (please), kalo for 'kalau' (Indonesian language, meaning 'if' in English) Mixing word and numeric or special characters [3,7,27,39,50] ri8 (right), 2morrow (tomorrow), ni8t (night), orang2 (Indonesian language, meaning people in English) Word exaggeration [3,7,27,39,45,47,[49][50][51]64] goood (good), Pleasssseee (please), cooool (cool), helloooo (hello) Abbreviated words [3,39,45,49,56,64] bght (brought or bought), tkt (ticket), flm (film), TC (take care)…”
Section: Type Of Non-standard Word Example
mentioning
confidence: 99%
“…Austronesian & Germanic Indonesian-English [7] Malay-English [57] Dravidian & Germanic Kannada-English [44] Malayalam-English [47,64] Tamil-English [47,51] Telugu-English [50] Germanic & Trans Eurasian Dutch-Turkish [46] Germanic & Germanic Dutch-English [20,46,49] Dutch-Limburgish [43] German-English [46] Indo-Aryan & Germanic Assamese-English [21] Assamese-Hindi-Bengali-English [18,63] Bengali-English [11,27,54] Bengali-Hindi-English [54] Gujarati-English [59] Gujarati-Hindi-English [58] Hindi-English [11,15,37,48,52,54,60,61,65] Konkani-English [45] Punjabi-English [16] Sinhala-English [3,56] Italic & Germanic French-English [46] French-Italian-Spanish-English [55] Portuguese-English [46] Spanish-English [15,38,41,46,49] Italic & American Spanish-Wixarika…”
Section: Language Family Combination
mentioning
confidence: 99%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Singh et al [20] build an automatic named entity recognition (NER) system for Hindi-English code-mixed data and the proposed system outperformed with 33.18% of F1-score in comparison with existing baseline systems. However, this work can be extended to build natural language processing (NLP) [21] presented a supervised learning model for word-level language identification in Bengali-English code-mixed data. Two types of word encoding methods such as character and phonetic are used along with stacking and threshold techniques.…”
Section: Issn: 2088-8708
mentioning
confidence: 99%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Another RNN variation technique, LSTM, has shown satisfactory performance in identifying Hindi-English and Bengali-English code-mixed text [27,52,54]. In [52], the LSTM architecture could give a high average F1 score of 93.4% and an average accuracy of 96.1% across the three classes.…”
Section: 1) Machine Learning Approach
mentioning
confidence: 99%
“…The following we identified some non-standard words encountered from the investigated papers. We categorised the non-standard words into four types, such as non-standard spelling [7,15,56], abbreviated words [3,37,39,45,49,56,64], exaggerated words [3,7,27,39,45,47,[49][50][51]64], and mixing characters with numbers or special characters [3,27,39,50]. Table 6 describes some examples of non-standard words found in code-mixed text LID.…”
Section: ) Non-standard Words
mentioning
confidence: 99%
“…Non-standard spelling [7,56] Prends or prenzz (friends), plis (please), kalo for 'kalau' (Indonesian language, meaning 'if' in English) Mixing word and numeric or special characters [3,7,27,39,50] ri8 (right), 2morrow (tomorrow), ni8t (night), orang2 (Indonesian language, meaning people in English) Word exaggeration [3,7,27,39,45,47,[49][50][51]64] goood (good), Pleasssseee (please), cooool (cool), helloooo (hello) Abbreviated words [3,39,45,49,56,64] bght (brought or bought), tkt (ticket), flm (film), TC (take care)…”
Section: Type Of Non-standard Word Example
mentioning
confidence: 99%
“…Austronesian & Germanic Indonesian-English [7] Malay-English [57] Dravidian & Germanic Kannada-English [44] Malayalam-English [47,64] Tamil-English [47,51] Telugu-English [50] Germanic & Trans Eurasian Dutch-Turkish [46] Germanic & Germanic Dutch-English [20,46,49] Dutch-Limburgish [43] German-English [46] Indo-Aryan & Germanic Assamese-English [21] Assamese-Hindi-Bengali-English [18,63] Bengali-English [11,27,54] Bengali-Hindi-English [54] Gujarati-English [59] Gujarati-Hindi-English [58] Hindi-English [11,15,37,48,52,54,60,61,65] Konkani-English [45] Punjabi-English [16] Sinhala-English [3,56] Italic & Germanic French-English [46] French-Italian-Spanish-English [55] Portuguese-English [46] Spanish-English [15,38,41,46,49] Italic & American Spanish-Wixarika…”
Section: Language Family Combination
mentioning
confidence: 99%