NeuMorph

Chakrabarty, Abhisek; Chaturvedi, Akshay; Garain, Utpal

doi:10.1145/3342354

Cited by 1 publication

(2 citation statements)

References 15 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…On the other, the scarcity of training data is not only problematic per se, but also because of its impact on the rest of trials. So, generating high-quality vector representations remains a challenge [26] and the imbalance in the training samples that start the long tail and bias phenomena is more likely, together with the proneness to overfitting of DL models [27], which can result in poor predictive power, thereby compromising both inference and decision making.…”

Section: Introductionmentioning

confidence: 99%

“…These encompass a class of NLP problems that involve the assignment of a categorical label to each member of a sequence of observed values, and whose output facilitates downstream applications, such as parsing or semantic analysis, so errors at this stage can lower their performance [31]. Among the most important, we can highlight named entity recognition [7,32], multi-word expression identification [29], and morphological [26] and POS tagging [2,[33][34][35][36]. It is precisely in this framework, the generation of POS taggers for low-resource scenarios by means of non-deep ML, that we propose the study of model selection based on the early estimation of learning curves.…”

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

Surfing the Modeling of pos Taggers in Low-Resource Scenarios

et al. 2022

View full text Add to dashboard Cite

The recent trend toward the application of deep structured techniques has revealed the limits of huge models in natural language processing. This has reawakened the interest in traditional machine learning algorithms, which have proved still to be competitive in certain contexts, particularly in low-resource settings. In parallel, model selection has become an essential task to boost performance at reasonable cost, even more so when we talk about processes involving domains where the training and/or computational resources are scarce. Against this backdrop, we evaluate the early estimation of learning curves as a practical mechanism for selecting the most appropriate model in scenarios characterized by the use of non-deep learners in resource-lean settings. On the basis of a formal approximation model previously evaluated under conditions of wide availability of training and validation resources, we study the reliability of such an approach in a different and much more demanding operational environment. Using as a case study the generation of pos taggers for Galician, a language belonging to the Western Ibero-Romance group, the experimental results are consistent with our expectations.

show abstract

Section: Introductionmentioning

confidence: 99%