Noha Y. Hassan scite author profile

With the evolution of social media platforms, the Internet is used as a source for obtaining news about current events. Recently, Twitter has become one of the most popular social media platforms that allows public users to share the news. The platform is growing rapidly especially among young people who may be influenced by the information from anonymous sources. Therefore, predicting the credibility of news in Twitter becomes a necessity especially in the case of emergencies. This paper introduces a classification model based on supervised machine learning techniques and word-based N-gram analysis to classify Twitter messages automatically into credible and not credible. Five different supervised classification techniques are applied and compared namely: Linear Support Vector Machines (LSVM), Logistic Regression (LR), Random Forests (RF), Naïve Bayes (NB) and K-Nearest Neighbors (KNN). The research investigates two feature representations (TF and TF-IDF) and different word N-gram ranges. For model training and testing, 10-fold cross validation is performed on two datasets in different languages (English and Arabic). The best performance is achieved using a combination of both unigrams and bigrams, LSVM as a classifier and TF-IDF as a feature extraction technique. The proposed model achieves 84.9% Accuracy, 86.6% Precision, 91.9% Recall, and 89% F-Measure on the English dataset. Regarding the Arabic dataset, the model achieves 73.2% Accuracy, 76.4% Precision, 80.7% Recall, and 78.5% F-Measure. The obtained results indicate that word N-gram features are more relevant for the credibility prediction compared with content and source-based features, also compared with character N-gram features. Experiments also show that the proposed model achieved an improvement when compared to two models existing in the literature.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Noha Y. Hassan

Supervised Learning Approach for Twitter Credibility Detection

Credibility Detection in Twitter Using Word N-gram Analysis and Supervised Machine Learning Techniques

Contact Info

Product

Resources

About