Profiling of OCR'ed Historical Texts Revisited

Fink, Florian; Schulz, Klaus U.; Springmann, Uwe

doi:10.1145/3078081.3078096

Cited by 12 publications

(9 citation statements)

References 4 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…In the revisited version of this method, Fink et al [49] additionally refine profiles from user feedback of manual correction steps. They also enlarge their set of patterns with addition ones found in documents of earlier periods.…”

Section: Isolated-word Approachesmentioning

confidence: 99%

Survey of Post-OCR Processing Approaches

et al. 2021

View full text Add to dashboard Cite

Optical character recognition (OCR) is one of the most popular techniques used for converting printed documents into machine-readable ones. While OCR engines can do well with modern text, their performance is unfortunately significantly reduced on historical materials. Additionally, many texts have already been processed by various out-of-date digitisation techniques. As a consequence, digitised texts are noisy and need to be post-corrected. This article clarifies the importance of enhancing quality of OCR results by studying their effects on information retrieval and natural language processing applications. We then define the post-OCR processing problem, illustrate its typical pipeline, and review the state-of-the-art post-OCR processing approaches. Evaluation metrics, accessible datasets, language resources, and useful toolkits are also reported. Furthermore, the work identifies the current trend and outlines some research directions of this field.

show abstract

Section: Isolated-word Approachesmentioning

confidence: 99%

Survey of Post-OCR Processing Approaches

et al. 2021

View full text Add to dashboard Cite

show abstract

“…Therefore, we aim to include a trainable pixel classifier in order to either provide a valid starting point for other segmentation approaches by classifying pixels and consequently connected contours as text, image, and noise or even perform a fine-grained semantic markup [46]. Of course, a more powerful segmentation approach must also comprise a more sophisticated method for the determination of the reading order which 39 https://www.cost.eu/cost-actions/ 40 https://www.distant-reading.net/ also has to be integrated into LAREX. To generate the reading order, the idea is to allow the user to comfortably specify rules based on the detected region types as well as their absolute and relative position.…”

Section: Future Workmentioning

confidence: 99%

OCR4all - An Open-Source Tool Providing a (Semi-)Automatic OCR Workflow for Historical Printings

Reul¹,

Christ²,

Hartelt³

et al. 2019

Preprint

Self Cite

View full text Add to dashboard Cite

Optical Character Recognition (OCR) on historical printings is a challenging task mainly due to the complexity of the layout and the highly variant typography. Nevertheless, in the last few years great progress has been made in the area of historical OCR, resulting in several powerful open-source tools for preprocessing, layout recognition and segmentation, character recognition and post-processing. The drawback of these tools often is their limited applicability by non-technical users like humanist scholars and in particular the combined use of several tools in a workflow. In this paper we present an open-source OCR software called OCR4all, which combines state-of-the-art OCR components and continuous model training into a comprehensive workflow. A comfortable GUI allows error corrections not only in the final output, but already in early stages to minimize error propagations. Further on, extensive configuration capabilities are provided to set the degree of automation of the workflow and to make adaptations to the carefully selected default parameters for specific printings, if necessary. Experiments showed that users with minimal or no experience were able to capture the text of even the earliest printed books with manageable effort and great quality, achieving excellent character error rates (CERs) below 0.5%. The fully automated application on 19th century novels showed that OCR4all can considerably outperform the commercial state-of-the-art tool ABBYY Finereader on moderate layouts if suitably pretrained mixed OCR models are available. The architecture of OCR4all allows the easy integration (or substitution) of newly developed tools for its main components by standardized interfaces like PageXML, thus aiming at continual higher automation for historical printings.

show abstract

“…The system is under active development which resulted in several improvements on the original approach. In [40] Fink et al added three major extensions: First, making the system more adaptive to manual interventions of the user increased the precision with respect to identifying erroneous OCR tokens. Second, the linguistic background resources were extended by new historical patterns which leads a more successful discrimination of historical spelling from real OCR errors.…”

Section: Pocotomentioning

confidence: 99%

OCR4all—An Open-Source Tool Providing a (Semi-)Automatic OCR Workflow for Historical Printings

et al. 2019

Self Cite

View full text Add to dashboard Cite

Optical Character Recognition (OCR) on historical printings is a challenging task mainly due to the complexity of the layout and the highly variant typography. Nevertheless, in the last few years great progress has been made in the area of historical OCR, resulting in several powerful open-source tools for preprocessing, layout recognition and segmentation, character recognition and post-processing. The drawback of these tools often is their limited applicability by non-technical users like humanist scholars and in particular the combined use of several tools in a workflow. In this paper we present an open-source OCR software called OCR4all, which combines state-of-the-art OCR components and continuous model training into a comprehensive workflow. A comfortable GUI allows error corrections not only in the final output, but already in early stages to minimize error propagations. Further on, extensive configuration capabilities are provided to set the degree of automation of the workflow and to make adaptations to the carefully selected default parameters for specific printings, if necessary. Experiments showed that users with minimal or no experience were able to capture the text of even the earliest printed books with manageable effort and great quality, achieving excellent character error rates (CERs) below 0.5%. The fully automated application on 19 th century novels showed that OCR4all can considerably outperform the commercial state-of-the-art tool ABBYY Finereader on moderate layouts if suitably pretrained mixed OCR models are available. The architecture of OCR4all allows the easy integration (or substitution) of newly developed tools for its main components by standardized interfaces like PageXML, thus aiming at continual higher automation for historical printings.

show abstract

Profiling of OCR'ed Historical Texts Revisited

Cited by 12 publications

References 4 publications

Survey of Post-OCR Processing Approaches

Survey of Post-OCR Processing Approaches

OCR4all - An Open-Source Tool Providing a (Semi-)Automatic OCR Workflow for Historical Printings

OCR4all—An Open-Source Tool Providing a (Semi-)Automatic OCR Workflow for Historical Printings

Contact Info

Product

Resources

About