A prototype gutenberg-hathitrust sentence-level parallel corpus for OCR error analysis

Ji, Meng; Dubnicek, Ryan; Worthey, Glen; Underwood, Ted; Downie, J. Stephen

doi:10.1145/3529372.3533298

Cited by 2 publications

(2 citation statements)

References 18 publications

(21 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…With the increasing popularity of applying NLP techniques to DL textual resources for macro-level computation research [24,4], concerns about the reliability of NLP techniques for processing digitized library collections have recently been on the rise [27,45,48]. Based on our literature review, one of the major issues that challenges NLP techniques' reliability on OCR'd texts is their potential inclusion of errors resulting from the OCR process [27,45,48].…”

Section: Impact Of Ocr Errors On Downstream Nlp Tasksmentioning

confidence: 99%

“…Overall, existing work concentrating on the investigation of the impact of OCR errors on NLP tasks can be divided into two groups. One is based on quantitative analysis [27,45]. In this group, researchers usually measure and compare the performance differences of the same NLP tool applied on the clean versus the OCR'd version of texts.…”

Section: Impact Of Ocr Errors On Downstream Nlp Tasksmentioning

confidence: 99%

See 1 more Smart Citation

Evaluating BERT-based scientific relation classifiers for scholarly knowledge graph construction on digital library collections

D’Souza

Auer

et al. 2021

Int J Digit Libr

View full text Add to dashboard Cite

The rapid growth of research publications has placed great demands on digital libraries (DL) for advanced information management technologies. To cater to these demands, techniques relying on knowledge-graph structures are being advocated. In such graph-based pipelines, inferring semantic relations between related scientific concepts is a crucial step. Recently, BERT-based pre-trained models have been popularly explored for automatic relation classification. Despite significant progress, most of them were evaluated in different scenarios, which limits their comparability. Furthermore, existing methods are primarily evaluated on clean texts, which ignores the digitization context of early scholarly publications in terms of machine scanning and optical character recognition (OCR). In such cases the

show abstract