Md. Shymon Islam scite author profile

Md. Shymon Islam

3Publications

1Citation Statement Received

33Citation Statements Given

How they've been cited

How they cite others

Affiliations

Khulna University

Publications

Order By: Most citations

Compilation, Analysis and Application of a Comprehensive Bangla Corpus KUMono

et al. 2022

View full text Add to dashboard Cite

Research in Natural Language Processing (NLP) and computational linguistics highly depends on a good quality representative corpus of any specific language. Bangla is one of the most spoken languages in the world but Bangla NLP research is in its early stage of development due to the lack of quality public corpus. This article describes the detailed compilation methodology of a comprehensive monolingual Bangla corpus, KUMono. The newly developed corpus consists of more than 350 million word tokens and more than one million unique tokens from 18 major text categories of online Bangla websites. We have conducted several word-level and character-level linguistic phenomenon analyses based on empirical studies of the developed corpus. The corpus follows Zipf's curve and hapax legomena rule. The quality of the corpus is also assessed by analyzing and comparing the inherent sparseness of the corpus with existing Bangla corpora, by analyzing the distribution of function words of the corpus and vocabulary growth rate. We have developed a Bangla article categorization application based on the KUMono corpus and received compelling results by comparing to the state-of-the-art models.

show abstract

Thyroid Disease Prediction based on Feature Selection and Machine Learning

Peya

Islam²,

Chumki³

2022

View full text Add to dashboard Cite

A solution method to maximal covering location problem based on chemical reaction optimization (CRO) algorithm

Islam

2023

Soft Comput

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Md. Shymon Islam

Compilation, Analysis and Application of a Comprehensive Bangla Corpus KUMono

Thyroid Disease Prediction based on Feature Selection and Machine Learning

A solution method to maximal covering location problem based on chemical reaction optimization (CRO) algorithm

Contact Info

Product

Resources

About