Jianqi Ma scite author profile

This paper introduces a novel rotation-based framework for arbitrary-oriented text detection in natural scene images. We present the Rotation Region Proposal Networks (RRPN), which are designed to generate inclined proposals with text orientation angle information. The angle information is then adapted for bounding box regression to make the proposals more accurately fit into the text region in terms of the orientation. The Rotation Region-of-Interest (RRoI) pooling layer is proposed to project arbitrary-oriented proposals to a feature map for a text region classifier. The whole framework is built upon a regionproposal-based architecture, which ensures the computational efficiency of the arbitrary-oriented text detection compared with previous text detection systems. We conduct experiments using the rotation-based framework on three real-world scene text detection datasets and demonstrate its superiority in terms of effectiveness and efficiency over previous approaches.

show abstract

Text Gestalt: Stroke-Aware Scene Text Image Super-resolution

Chen

et al. 2022

AAAI

View full text Add to dashboard Cite

In the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have been proposed to tackle this problem, they usually treat text images as general images while ignoring the fact that the visual quality of strokes (the atomic unit of text) plays an essential role for text recognition. According to Gestalt Psychology, humans are capable of composing parts of details into the most similar objects guided by prior knowledge. Likewise, when humans observe a low-resolution text image, they will inherently use partial stroke-level details to recover the appearance of holistic characters. Inspired by Gestalt Psychology, we put forward a Stroke-Aware Scene Text Image Super-Resolution method containing a Stroke-Focused Module (SFM) to concentrate on stroke-level internal structures of characters in text images. Specifically, we attempt to design rules for decomposing English characters and digits at stroke-level, then pre-train a text recognizer to provide stroke-level attention maps as positional clues with the purpose of controlling the consistency between the generated super-resolution image and high-resolution ground truth. The extensive experimental results validate that the proposed method can indeed generate more distinguishable images on TextZoom and manually constructed Chinese character dataset Degraded-IC13. Furthermore, since the proposed SFM is only used to provide stroke-level guidance when training, it will not bring any time overhead during the test phase. Code is available at https://github.com/FudanVI/FudanOCR/tree/main/text-gestalt.

show abstract

Text Prior Guided Scene Text Image Super-Resolution

Guo

Zhang

2023

IEEE Trans. on Image Process.

View full text Add to dashboard Cite

BTS: A Bi-lingual Benchmark for Text Segmentation in the Wild

et al. 2022

View full text Add to dashboard Cite

Face Recognition via Active Annotation and Learning

Hao

Wang

et al. 2016

View full text Add to dashboard Cite

A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolution

Liang²,

Zhang

2022

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Jianqi Ma

Arbitrary-Oriented Scene Text Detection via Rotation Proposals

Text Gestalt: Stroke-Aware Scene Text Image Super-resolution

Text Prior Guided Scene Text Image Super-Resolution

BTS: A Bi-lingual Benchmark for Text Segmentation in the Wild

Face Recognition via Active Annotation and Learning

A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolution

Contact Info

Product

Resources

About