Research Paper:
VPDL: Visual Prompt-Guided Differential Learning for Generalizable Scene Text Recognition
Yali Niu*, Jiahao An*,, and Shihao Zou**

*School of Transportation & Information, Shaanxi College of Communications Technology
No.19 Wenjing Road, Weiyang District, Xi’an, Shaanxi 710018, China
Corresponding author
**School of Computer Science and Technology, Huazhong University of Science and Technology
No.1037 Luoyu Road, Hongshan District, Wuhan, Hubei 430074, China
Scene text recognition (STR) in natural images remains highly challenging due to the large variations in character appearance across diverse real-world conditions, such as changes in font, color, layout, and background complexity—which hinder model generalization and remain insufficiently explored. To address this issue, we propose a visual prompt-guided differential learning (VPDL) framework designed to improve the generalization capability of STR models without requiring scene-specific fine-tuning. Inspired by the human ability to reference prior visual knowledge when recognizing text, VPDL introduces a set of character-level visual prompts that guide the model in perceiving appearance variations among characters. Built upon these prompts, we develop a local-to-global differential learning strategy that enhances patch-level representations and aligns global features with character cues while preserving scene-specific information. Additionally, to mitigate exposure bias in autoregressive decoding, we replace conventional label inputs with context-aware textual prompts, encouraging the decoder to better utilize textual cues embedded in image features. Extensive experiments on widely used benchmarks and real-world datasets demonstrate the effectiveness of VPDL.
Prompt-guided STR model
- [1] Z. Zou, K. Chen, Z. Shi, Y. Guo, and J. Ye, “Object Detection in 20 Years: A Survey,” Proc. of the IEEE, Vol.111, Issue 3, pp. 257-276, 2023. https://doi.org/10.1109/JPROC.2023.3238524
- [2] Y. Chen, X. Yuan, J. Wang, R. Wu, X. Li, Q. Hou, and M.-M. Cheng, “YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.47, Issue 6, pp. 4240-4252, 2025. https://doi.org/10.1109/TPAMI.2025.3538473
- [3] M. Fang, X. Rui, H. Cheng, X. Liu, J. She, Y. Du, and H. Tan, “Small Object Detection Algorithm Based on Improved Attention Mechanism and Feature Fusion of YOLOv8,” J. Adv. Comput. Intell. Intell. Inform., Vol.29, No.4, pp. 941-955, 2025. https://doi.org/10.20965/jaciii.2025.p0941
- [4] J. Wan, S. Song, W. Yu, Y. Liu, W. Cheng, F. Huang, X. Bai, C. Yao, and Z. Yang, “OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 15641-15653, 2024. https://doi.org/10.1109/CVPR52733.2024.01481
- [5] B. Zhang, H. Xie, Z. Gao, and Y. Wang, “Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 28358-28368, 2024. https://doi.org/10.1109/CVPR52733.2024.02679
- [6] Q. Jiang, J. Wang, D. Peng, C. Liu, and L. Jin, “Revisiting Scene Text Recognition: A Data Perspective,” 2023 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 20486-20497, 2023. https://doi.org/10.1109/ICCV51070.2023.01878
- [7] C. Luo, C. Cheng, Q. Zheng, and C. Yao, “GeoLayoutLM: Geometric Pre-training for Visual Information Extraction,” 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 7092-7101, 2023. https://doi.org/10.1109/CVPR52729.2023.00685
- [8] R. Wang, Y. Xue, and L. Jin, “DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations,” Proc. of the AAAI Conf. on Artificial Intelligence, Vol.38, No.6, pp. 5563-5571, 2024. https://doi.org/10.1609/aaai.v38i6.28366
- [9] M. Peng, K. Chen, X. Guo, Q. Zhang, H. Zhong, M. Zhu, and H. Yang, “Diffusion Models for Intelligent Transportation Systems: A Survey,” IEEE Trans. on Intelligent Transportation Systems, Vol.26, Issue 12, pp. 21526-21543, 2025. https://doi.org/10.1109/TITS.2025.3613178
- [10] M. Bakirci, “Advanced aerial monitoring and vehicle classification for intelligent transportation systems with YOLOv8 variants,” J. of Network and Computer Applications, Vol.237, Article No.104134, 2025. https://doi.org/10.1016/j.jnca.2025.104134
- [11] M. A. Jan, M. Adil, B. Brik, S. Harous, and S. Abbas, “Making Sense of Big Data in Intelligent Transportation Systems: Current Trends, Challenges and Future Directions,” ACM Comput. Surv., Vol.57, Issue 8, Article No.197, 2025. https://doi.org/10.1145/3716371
- [12] C. Zhang, Y. Tao, K. Du, W. Ding, B. Wang, J. Liu, and W. Wang, “Character-Level Street View Text Spotting Based on Deep Multisegmentation Network for Smarter Autonomous Driving,” IEEE Trans. on Artificial Intelligence, Vol.3, Issue 2, pp. 297-308, 2022. https://doi.org/10.1109/TAI.2021.3116216
- [13] L. Wu, C. Zhang, J. Liu, J. Han, J. Liu, E. Ding, and X. Bai, “Editing Text in the Wild,” Proc. of the 27th ACM Int. Conf. on Multimedia, pp. 1500-1508, 2019. https://doi.org/10.1145/3343031.3350929
- [14] B. Shi, X. Bai, and C. Yao, “An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.39, Issue 11, pp. 2298-2304, 2017. https://doi.org/10.1109/TPAMI.2016.2646371
- [15] S. Fang, H. Xie, Y. Wang, Z. Mao, and Y. Zhang, “Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 7094-7103, 2021. https://doi.org/10.1109/CVPR46437.2021.00702
- [16] D. Bautista and R. Atienza, “Scene Text Recognition with Permuted Autoregressive Sequence Models,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 178-196, 2022. https://doi.org/10.1007/978-3-031-19815-1_11
- [17] Y. Du, Z. Chen, C. Jia, X. Yin, T. Zheng, C. Li, Y. Du, and Y.-G. Jiang, “SVTR: Scene Text Recognition with a Single Visual Model,” arXiv preprint, arXiv:2205.00159, 2022. https://doi.org/10.48550/arXiv.2205.00159
- [18] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” Proc. of the 23rd Int. Conf. on Machine Learning, pp. 369-376, 2006. https://doi.org/10.1145/1143844.1143891
- [19] Z. Cheng, F. Bai, Y. Xu, G. Zheng, S. Pu, and S. Zhou, “Focusing Attention: Towards Accurate Text Recognition in Natural Images,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 5086-5094, 2017. https://doi.org/10.1109/ICCV.2017.543
- [20] F. Xue, J. Sun, Y. Xue, Q. Wu, L. Zhu, X. Chang, and S.-C. Cheung, “Attention Guidance by Cross-Domain Supervision Signals for Scene Text Recognition,” IEEE Trans. on Image Processing, Vol.34, pp. 717-728, 2025. https://doi.org/10.1109/TIP.2024.3523799
- [21] J. Lee, S. Park, J. Baek, S. J. Oh, S. Kim, and H. Lee, “On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 2326-2335, 2020. https://doi.org/10.1109/CVPRW50498.2020.00281
- [22] J. Xu, Y. Wang, H. Xie, and Y. Zhang, “OTE: Exploring Accurate Scene Text Recognition Using One Token,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 28327-28336, 2024. https://doi.org/10.1109/CVPR52733.2024.02676
- [23] W. Hu, X. Cai, J. Hou, S. Yi, and Z. Lin, “GTC: Guided Training of CTC Towards Efficient and Accurate Scene Text Recognition,” Proc. of the AAAI Conf. on Artificial Intelligence, Vol.34, No.7, pp. 11005-11012, 2020. https://doi.org/10.1609/aaai.v34i07.6735
- [24] M. Yousef and T. E. Bishop, “OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by Learning to Unfold,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 14698-14707, 2020. https://doi.org/10.1109/CVPR42600.2020.01472
- [25] T. Wang, Y. Zhu, L. Jin, D. Peng, Z. Li, M. He, Y. Wang, and C. Luo, “Implicit Feature Alignment: Learn to Convert Text Recognizer to Text Spotter,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 5969-5978, 2021. https://doi.org/10.1109/CVPR46437.2021.00591
- [26] Y. Du, C. Li, R. Guo, X. Yin, W. Liu, J. Zhou, Y. Bai, Z. Yu, Y. Yang, Q. Dang, and H. Wang, “PP-OCR: A Practical Ultra Lightweight OCR System,” arXiv preprint, arXiv:2009.09941, 2020. https://doi.org/10.48550/arXiv.2009.09941
- [27] P. Wang, C. Da, and C. Yao, “Multi-granularity Prediction for Scene Text Recognition,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 339-355, 2022. https://doi.org/10.1007/978-3-031-19815-1_20
- [28] Y. Wang, H. Xie, S. Fang, J. Wang, S. Zhu, and Y. Zhang, “From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network,” 2021 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 14174-14183, 2021. https://doi.org/10.1109/ICCV48922.2021.01393
- [29] Z. Wan, M. He, H. Chen, X. Bai, and C. Yao, “TextScanner: Reading Characters in Order for Robust Scene Text Recognition,” Proc. of the AAAI Conf. on Artificial Intelligence, pp. 12120-12127, 2020. https://doi.org/10.1609/aaai.v34i07.6891
- [30] D. Yu, X. Li, C. Zhang, T. Liu, J. Han, J. Liu, and E. Ding, “Towards Accurate Scene Text Recognition with Semantic Reasoning Networks,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 12110-12119, 2020. https://doi.org/10.1109/CVPR42600.2020.01213
- [31] Z. Zhao, J. Tang, C. Lin, B. Wu, C. Huang, H. Liu, X. Tan, Z. Zhang, and Y. Xie, “Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 15567-15576, 2024. https://doi.org/10.1109/CVPR52733.2024.01474
- [32] B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “ASTER: An Attentional Scene Text Recognizer with Flexible Rectification,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.41, Issue 9, pp. 2035-2048, 2019. https://doi.org/10.1109/TPAMI.2018.2848939
- [33] Z. Cheng, F. Bai, Y. Xu, G. Zheng, S. Pu, and S. Zhou, “Focusing Attention: Towards Accurate Text Recognition in Natural Images,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 5086-5094, 2017. https://doi.org/10.1109/ICCV.2017.543
- [34] J. Baek, G. Kim, J. Lee, S. Park, D. Han, S. Yun, S. J. Oh, and H. Lee, “What is Wrong with Scene Text Recognition Model Comparisons? Dataset and Model Analysis,” 2019 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 4714-4722, 2019. https://doi.org/10.1109/ICCV.2019.00481
- [35] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Proc. of the Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol.1, pp. 4171-4186, 2019. https://doi.org/10.18653/v1/N19-1423
- [36] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” Proc. of the 34th Int. Conf. on Neural Information Processing Systems (NIPS’20), pp. 1877-1901, 2020. https://doi.org/10.18653/v1/N19-1423
- [37] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” OpenAI Blog, 2019.
- [38] A. Bar, Y. Gandelsman, T. Darrell, A. Globerson, and A. A. Efros, “Visual Prompting via Image Inpainting,” Advances in Neural Information Processing Systems, pp. 25005-25017, 2022.
- [39] J. Zhu and G. Pang, “Toward Generalist Anomaly Detection via In-Context Residual Learning with Few-Shot Sample Prompts,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 17826-17836, 2024. https://doi.org/10.1109/CVPR52733.2024.01688
- [40] F. Li, Q. Jiang, H. Zhang, T. Ren, S. Liu, X. Zou, H. Xu, H. Li, J. Yang, C. Li, L. Zhang, and J. Gao, “Visual in-Context Prompting,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 12861-12871, 2024. https://doi.org/10.1109/CVPR52733.2024.01222
- [41] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16times16 Words: Transformers for Image Recognition at Scale,” arXiv preprint, arXiv:2010.11929, 2020. https://doi.org/10.48550/arXiv.2010.11929
- [42] A. Mishra, K. Alahari, and C. Jawahar, “Scene Text Recognition using Higher Order Language Priors,” British Machine Vision Conf. (BMVC2), pp. 127.1-127.11, 2012. https://doi.org/10.5244/C.26.127
- [43] D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. Gomez i Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. A. Almazàn, and L.-P. de las Heras, “ICDAR 2013 Robust Reading Competition,” 2013 12th Int. Conf. on Document Analysis and Recognition, pp. 1484-1493, 2013. https://doi.org/10.1109/ICDAR.2013.221
- [44] D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny, “ICDAR 2015 Competition on Robust Reading,” 2015 13th Int. Conf. on Document Analysis and Recognition (ICDAR), pp. 1156-1160, 2015. https://doi.org/10.1109/ICDAR.2015.7333942
- [45] K. Wang, B. Babenko, and S. Belongie, “End-to-end scene text recognition,” 2011 IEEE Int. Conf. on Computer Vision, pp. 1457-1464, 2011. https://doi.org/10.1109/ICCV.2011.6126402
- [46] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing Text with Perspective Distortion in Natural Scenes,” 2013 IEEE Int. Conf. on Computer Vision, pp. 569-576, 2013. https://doi.org/10.1109/ICCV.2013.76
- [47] A. Risnumawan, P. Shivakumara, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” Expert Systems with Applications, Vol.41, Issue 18, pp. 8027-8048, 2014. https://doi.org/10.1016/j.eswa.2014.07.008
- [48] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, “Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 303-321, 2022. https://doi.org/10.1007/978-3-031-19815-1_18
- [49] A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie, “COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images,” arXiv preprint, arXiv:1601.07140, 2016. https://doi.org/10.48550/arXiv.1601.07140
- [50] C. K. Chng, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding et al., “ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text – RRC-ArT,” 2019 Int. Conf. on Document Analysis and Recognition (ICDAR), pp. 1571-1576, 2019. https://doi.org/10.1109/ICDAR.2019.00252
- [51] Y. Zhang, L. Gueguen, I. Zharkov, P. Zhang, K. Seifert, and B. Kadlec, “Uber-text: A large-scale dataset for optical character recognition from street-level imagery,” SUNw: Scene Understanding Workshop-CVPR, Vol.2017, 2017.
- [52] F. Sheng, Z. Chen, and B. Xu, “NRTR: A No-Recurrence Sequence-to-Sequence Model for Scene Text Recognition,” 2019 Int. Conf. on Document Analysis and Recognition, pp. 781-786, 2019. https://doi.org/10.1109/ICDAR.2019.00130
- [53] X. Yue, Z. Kuang, C. Lin, H. Sun, and W. Zhang, “RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition,” A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm (Eds.), “Computer Vision – ECCV 2020,” Lecture Notes in Computer Science, Vol.12364, pp. 135-151, 2020. https://doi.org/10.1007/978-3-030-58529-7_9
- [54] B. Na, Y. Kim, and S. Park, “Multi-modal Text Recognition Networks: Interactive Enhancements Between Visual and Semantic Features,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 446-463, 2022. https://doi.org/10.1007/978-3-031-19815-1_26
- [55] W. Xu, E. Eli, A. Aysa, X. Xu, and K. Ubul, “ASANet: Scene Text Recognition with Alternate Self-Attention,” 2025 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP), 2025. https://doi.org/10.1109/ICASSP49660.2025.10889330
- [56] C. Da, P. Wang, and C. Yao, “Multi-granularity prediction with learnable fusion for scene text recognition,” Int. J. of Computer Vision, Vol.134, No.1, Article No.47, 2026. https://doi.org/10.1007/s11263-025-02653-7
- [57] S. Zhao, R. Quan, L. Zhu, and Y. Yang, “CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-Trained Vision-Language Model,” IEEE Trans. on Image Processing, Vol.33, pp. 6893-6904, 2024. https://doi.org/10.1109/TIP.2024.3512354
- [58] R. Atienza, “Vision Transformer for Fast and Efficient Scene Text Recognition,” J. Lladós, D. Lopresti, and S. Uchida, “Document Analysis and Recognition – ICDAR 2021,” Lecture Notes in Computer Science, Vol.12821, pp. 319-334, 2021. https://doi.org/10.1007/978-3-030-86549-8_21
- [59] J. Baek, Y. Matsui, and K. Aizawa, “What If We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition with Fewer Labels,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 3112-3121, 2021. https://doi.org/10.1109/CVPR46437.2021.00313
- [60] M. Yang, M. Liao, P. Lu, J. Wang, S. Zhu, H. Luo, Q. Tian, and X. Bai, “Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text Recognition,” Proc. of the 30th ACM Int. Conf. on Multimedia (MM’22), pp. 4214-4223, 2022. https://doi.org/10.1145/3503161.3547784
- [61] X. Yang, Z. Qiao, J. Wei, D. Yang, and Y. Zhou, “Masked and Permuted Implicit Context Learning for Scene Text Recognition,” IEEE Signal Processing Letters, Vol.31, pp. 964-968, 2024. https://doi.org/10.1109/LSP.2024.3381893
- [62] Z. Wang, H. Xie, Y. Wang, J. Xu, B. Zhang, and Y. Zhang, “Symmetrical Linguistic Feature Distillation with CLIP for Scene Text Recognition,” Proc. of the 31st ACM Int. Conf. on Multimedia (MM’23), pp. 509-518, 2023. https://doi.org/10.1145/3581783.3611769
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.