Scene Uyghur Recognition Based on Visual Prediction Enhancement

Liu, Yaqi; Kong, Fanjie; Xu, Miaomiao; Silamu, Wushour; Li, Yanbing

doi:10.3390/s23208610

Cited by 1 publication

(1 citation statement)

References 42 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Scene Text Recognition Methods Employing Joint Learning of Internal Language Models: Unlike external guided language models, internal language models typically participate in the entire model training process and learn contextual semantic information during training. Methods based on RNN [1,38,[46][47][48][49][50] capture the temporal information of characters in the text through RNN, thereby understanding the dependency relationships between characters in the text. The TRBA [1] model adopts BiLSTM for sequence modeling, aiming to better comprehend the contextual information between characters.…”

Section: Related Workmentioning

confidence: 99%

Collaborative Encoding Method for Scene Text Recognition in Low Linguistic Resources: The Uyghur Language Case Study

Xu,

Zhang,

et al. 2024

Applied Sciences

Self Cite

View full text Add to dashboard Cite

Current research on scene text recognition primarily focuses on languages with abundant linguistic resources, such as English and Chinese. In contrast, there is relatively limited research dedicated to low-resource languages. Advanced methods for scene text recognition often employ Transformer-based architectures. However, the performance of Transformer architectures is suboptimal when dealing with low-resource datasets. This paper proposes a Collaborative Encoding Method for Scene Text Recognition in the low-resource Uyghur language. The encoding framework comprises three main modules: the Filter module, the Dual-Branch Feature Extraction module, and the Dynamic Fusion module. The Filter module, consisting of a series of upsampling and downsampling operations, performs coarse-grained filtering on input images to reduce the impact of scene noise on the model, thereby obtaining more accurate feature information. The Dual-Branch Feature Extraction module adopts a parallel structure combining Transformer encoding and Convolutional Neural Network (CNN) encoding to capture local and global information. The Dynamic Fusion module employs an attention mechanism to dynamically merge the feature information obtained from the Transformer and CNN branches. To address the scarcity of real data for natural scene Uyghur text recognition, this paper conducted two rounds of data augmentation on a dataset of 7267 real images, resulting in 254,345 and 3,052,140 scene images, respectively. This process partially mitigated the issue of insufficient Uyghur language data, making low-resource scene text recognition research feasible. Experimental results demonstrate that the proposed collaborative encoding approach achieves outstanding performance. Compared to baseline methods, our collaborative encoding approach improves accuracy by 14.1%.

show abstract