Jolanta Kovalevskaitė scite author profile

Jolanta Kovalevskaitė

5Publications

3Citation Statements Received

9Citation Statements Given

How they've been cited

How they cite others

Affiliations

Vytautas Magnus University

Publications

Order By: Most citations

Lithuanian Pedagogic Corpus: Correlations Between Linguistic Features and Text Complexity

Boizou

Kovalevskaitė

Rimkutė

2020

View full text Add to dashboard Cite

This paper discusses the problem of automatic CEFR (CEFR – Common European Framework of Reference for Languages: https://www.coe.int/en/web/common-european-framework-reference-languages/level-descriptions.) level assignment to texts. We address the correlations between the lexical, morphological and syntactic features and the different CEFR levels of the texts in the Lithuanian Pedagogic Corpus. Only the texts from coursebooks showed the correlation of investigated linguistic features with text complexity. In the coursebook sub-part of the corpus, we observed that higher language proficiency levels are associated with more complex linguistic features: their number increases in texts of higher CEFR levels from A1 to B2 (e.g., non-finite verb forms, participles, adverbial participles and half participles, dative and instrumental noun cases or longer sentences).

show abstract

Light Verb Constructions in Lithuanian: Identification and Classification

Kovalevskaitė¹,

Rimkutė²

2020

StALan

View full text Add to dashboard Cite

Light verb constructions (LVCs) are verb-noun constructions in which the noun carries the semantic meaning and the verb is semantically reduced, when compared with its main meaning, for example, atlikti analizę (‘to perform an analysis’). LVCs in Lithuanian have not been addressed much so far. The analysis of Lithuanian LVCs was carried out as a part of the PARSEME project on verbal identification of multiword expressions (MWE). This paper aims at presenting some initial findings on the identification of LVCs in Lithuanian, based on the 1st edition of the PARSEME shared-task results (2017). We describe the identification process according to the semantic and syntactic features of LVCs (PARSEME guidelines 1.0 2017) and discuss the grammatical features of the identified Lithuanian LVCs. LVCs seem to be less frequent in Lithuanian than in other languages: they make up about 0.2% (215 instances) of the analysed 200,000 token corpus. Based on the number of different LVCs, there seem to be two groups of verbs functioning as light verbs: a relatively small group of common light verbs used in the most prototypical examples of Lithuanian LVCs (e.g., vykdyti ‘to perform’, atlikti ‘to perform’, daryti ‘to do’, and turėti ‘to have’) and a larger group of less common light verbs. Most of the nouns in analysed LVCs have suffixes -imas and -ymas, which are the most typical Lithuanian suffixes for deriving a noun from a verb. Almost 40% of all LVCs are used with 1–3 words intervening between a verb and a noun.

show abstract

Pedagogic Corpus of Lithuanian: A New Resource for Learning and Teaching Lithuanian as a Foreign Language

Kovalevskaitė

Rimkutė

2020

View full text Add to dashboard Cite

SummaryThe paper aims to present the first pedagogic corpus of Lithuanian i.e. monolingual specialized corpus, prepared for learning and teaching Lithuanian in a foreign language classroom. The corpus has been collected as a result of the project “Lithuanian Academic Scheme for International Cooperation in Baltic Studies”. It is motivated by the need to have a more appropriate resource which could be representative, authentic and relevant enough concerning the process of learning and teaching Lithuanian as it is known that language represented in other existing corpora of Lithuanian (e.g. Corpus of Contemporary Lithuanian, 140 m tokens) is too complex to use for learning activities. The pedagogic corpus includes authentic Lithuanian texts, selected using such criteria as a learner-relevant communicative function and genre. Spoken language as well as written language are represented in the corpus. The size of the corpus is 669.000 tokens: 111.000 tokens from texts and spoken language for A1–A2 levels, 558.000 tokens from texts and spoken language for B1–B2 levels (according to the CEFR – Common European Framework of Reference for Languages). In this paper, we aim to discuss in detail the written subpart of the corpus (containing 620.000 tokens) which includes levelled texts from coursebooks and unlevelled texts from other sources. The level-appropriate labels were assigned automatically to the texts from other sources and this text classification procedure is presented in the paper. The texts from coursebooks and other sources could be classified into 29 text types (dialogs, narratives, information, etc.) and 4 groups according to the communicative aims: informational texts, educational texts, advertising and fiction. Informational texts comprise the biggest part of the corpus; three mostly represented text types differ in coursebook texts and other sources: the most common coursebook texts are informational, narratives, and dialogs (appr. 78% of all coursebook texts). Texts from other sources are represented with richer diversity – appr. 73% of all texts from this subpart can be classified into 5 text types: subtitles, informational texts, educational texts, fiction, and advisory texts. The future work making pedagogic corpus available for learners and its possible application are presented in the closing remarks.

show abstract

Lietuvių kalbos daiktavardinių frazių žodyno vienaformiai pastovieji junginiai

Rimkutė¹,

Kovalevskaitė²

2015

View full text Add to dashboard Cite

Why do Infinite Forms Matter: Analysis of Verbs from the Lexical Database of Lithuanian Language Usage

Kovalevskaitė

Rimkutė

2023

View full text Add to dashboard Cite

From the corpus data, we observe that in the real language usage, the particular verb does not appear in all theoretically possible finite and infinite verb forms in the morphologically rich Lithuanian but is used in those forms which are relevant for the verb patterning. On the one hand, by teaching vocabulary, is it important to represent lexis in these relevant forms – frequently used forms, and, on the other hand, in grammar teaching, there is a need to provide learners with appropriate vocabulary, e.g., by teaching infinite forms, to use verbs, in the usage of which, these forms are relevant and frequent.In this paper, we provide language teaching practitioners with the data about the frequently used Lithuanian verbs and show which of them and how often appear in infinite forms (participles in passive and active voice, adverbial participles, half participles). As a research data we use 200 verbs from the Lexical Database of Lithuanian Language Usage which was developed on the basis of the written subcorpus of the Pedagogic corpus of Lithuanian. The investigated verbs belong to the frequent vocabulary: in the corpus of approx. 700,000 tokens, these verbs are used 100 times (and above). First, we analysed, which verbs appear in infinite forms, second, we checked whether frequent and typical infinite forms are included into corpus pattern(s) of these particular verbs, and if there is a link between the infinite form and a particular meaning of the verb.All verbs (except of three verbs with no infinite forms) were included into one of three groups: 1) 11 verbs which occur in the infinite forms frequently (more than 50% of all forms – finite and infinite) and, accordingly, typical; 2) 117 verbs with the infinite forms making up from 10 to 50%; 3) 69 verbs, with the infinite forms making up less than 10% of all verb forms. Interestingly, the verbs of the first group, usually have only one infinite form, e.g., participle in passive voice which makes up more than 50% of all forms of verb. These cases are also frequently observed in the second verb group. Thus, if the verb tends to be used in infinite forms, it is important to know which infinite form is relevant to that particular verb.In the Lexical Database of Lithuanian Language Usage, lexical and grammatical patterning of the word is represented in the form of corpus patterns. In this study, we showed the interrelation between the frequently used infinite forms of the verb and its corpus patterns (also, corpus patterns related to particular meaning of the polysemous verb). We can expect various applications of the provided data in the Lithuanian as a foreign language teaching: the provided data about the verbs typical and frequent in infinite forms and the corpus patterns including these infinite forms can be used for building vocabulary training as well as for developing grammar exercises.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.