Diachronic Treebanks for Historical Linguistics

Eckhoff, Hanne Martine; Luraghi, Silvia; Passarotti, Marco

doi:10.1075/bct.113

Benjamins Current Topics

2020

DOI: 10.1075/bct.113

|View full text |Cite

Diachronic Treebanks for Historical Linguistics

Hanne Martine Eckhoff¹,

Silvia Luraghi²,

Marco Passarotti³

Help me understand this report

Search citation statements

Order By: Relevance

Paper Sections

Select...

Citation Types

Supporting

Mentioning

Contrasting

Year Published

2023

2024

Publication Types

Select...

Article2

Relationship

Self Cite0

Independent2

Authors

Journals

Cited by 2 publications

References 4 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

Linguistic annotation of cuneiform texts using treebanks and deep learning

Ong,

Gordin

2024

Digital Scholarship in the Humanities

View full text Add to dashboard Cite

We describe an efficient pipeline for morpho-syntactically annotating an ancient language corpus which takes advantage of bootstrapping techniques. This pipeline is designed for ancient language scholars looking to jump-start their own treebank projects, which can in turn serve further pedagogical research projects in the target language. We situate our work in the field of similar ancient language treebank projects, arguing that our approach shows that individual humanities scholars can leverage current machine-learning tools to produce their own richly annotated corpora. We illustrate this pipeline by producing a new Akkadian-language treebank based on two volumes from the online editions of the State Archives of Assyria project hosted on Oracc, as well as a spaCy language model named AkkParser trained on that treebank. Both of these are made publicly available for annotating other Akkadian corpora. In addition, we discuss linguistic issues particular to the Neo-Assyrian letter corpus and data-encoding complications of cuneiform texts in Oracc. The strategies, language models, and processing scripts we developed to handle both linguistic and data-encoding issues in this project will be of special interest to scholars seeking to develop their own cuneiform treebanks.

show abstract

Linguistic annotation of cuneiform texts using treebanks and deep learning

Ong,

Gordin

2024

Digital Scholarship in the Humanities

View full text Add to dashboard Cite

show abstract

Reconstructing variation in Indo-European word order

et al. 2023

View full text Add to dashboard Cite

Word order is a central issue in the reconstruction of Proto-Indo-European syntax. Categorical approaches have proved to be inadequate because they postulate for the protolanguage a typological consistency which is absent in any of the attested daughter languages. Following recent research, we adopt a gradient approach to word order, which treats word order preferences as a continuous variable. We analyze four word order patterns based on data extracted from treebanks of ancient Indo-European languages. After presenting our results for AdpN/NAdp, GN/NG, AN/NA, and OV/VO, we draw a number of conclusions concerning variation within individual languages, crosslinguistic variation, and variation in diachrony that support the claim that variability should be taken as the normal state across languages, including reconstructed stages. We conclude that a non-discrete approach has the advantage of leading to a reconstruction that better conforms to the situation known from real languages, with variation as a key feature.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Diachronic Treebanks for Historical Linguistics

Cited by 2 publications

References 4 publications

Linguistic annotation of cuneiform texts using treebanks and deep learning

Linguistic annotation of cuneiform texts using treebanks and deep learning

Reconstructing variation in Indo-European word order

Contact Info

Product

Resources

About