Sahil Swami scite author profile

Sahil Swami

4Publications

42Citation Statements Received

36Citation Statements Given

How they've been cited

How they cite others

Affiliations

International Institute of Information Technology, Hyderabad

Publications

Order By: Most citations

Gender Prediction in English-Hindi Code-Mixed Social Media Content: Corpus and Baseline System

Khandelwal

Swami

Akhtar

et al. 2018

CyS

View full text Add to dashboard Cite

The tremendous amount of user generated data through social networking sites led to the gaining popularity of automatic text classification in the field of computational linguistics over the past decade. Within this domain, one problem that has drawn the attention of many researchers is automatic humor detection in texts. In depth semantic understanding of the text is required to detect humor which makes the problem difficult to automate. With increase in the number of social media users, many multilingual speakers often interchange between languages while posting on social media which is called code-mixing. It introduces some challenges in the field of linguistic analysis of social media content (Barman et al., 2014), like spelling variations and non-grammatical structures in a sentence. Past researches include detecting puns in texts (Kao et al., 2016) and humor in one-lines (Mihalcea et al., 2010) in a single language, but with the tremendous amount of code-mixed data available online, there is a need to develop techniques which detects humor in code-mixed tweets. In this paper, we analyze the task of humor detection in texts and describe a freely available corpus containing English-Hindi code-mixed tweets annotated with humorous(H) or non-humorous(N) tags. We also tagged the words in the tweets with Language tags (English/Hindi/Others). Moreover, we describe the experiments carried out on the corpus and provide a baseline classification system which distinguishes between humorous and non-humorous texts.

show abstract

A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection

Swami¹,

Khandelwal²,

Singh³

et al. 2018

Preprint

View full text Add to dashboard Cite

Social media platforms like twitter and facebook have become two of the largest mediums used by people to express their views towards different topics. Generation of such large user data has made NLP tasks like sentiment analysis and opinion mining much more important. Using sarcasm in texts on social media has become a popular trend lately. Using sarcasm reverses the meaning and polarity of what is implied by the text which poses challenge for many NLP tasks. The task of sarcasm detection in text is gaining more and more importance for both commercial and security services. We present the first English-Hindi code-mixed dataset of tweets marked for presence of sarcasm and irony where each token is also annotated with a language tag. We present a baseline supervised classification system developed using the same dataset which achieves an average F-score of 78.4 after using random forest classifier and performing 10-fold cross validation.

show abstract

Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System

Khandelwal¹,

Swami²,

Akhtar³

et al. 2018

Preprint

View full text Add to dashboard Cite

Gender Prediction in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System

Khandelwal¹,

Swami²,

Akhtar³

et al. 2018

Preprint

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Sahil Swami

Gender Prediction in English-Hindi Code-Mixed Social Media Content: Corpus and Baseline System

A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection

Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System

Gender Prediction in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System

Contact Info

Product

Resources

About