single-jc.php

JACIII Vol.30 No.4 pp. 1306-1317
(2026)

Research Paper:

VPDL: Visual Prompt-Guided Differential Learning for Generalizable Scene Text Recognition

Yali Niu*, Jiahao An*,†, and Shihao Zou** ORCID Icon

*School of Transportation & Information, Shaanxi College of Communications Technology
No.19 Wenjing Road, Weiyang District, Xi’an, Shaanxi 710018, China

Corresponding author

**School of Computer Science and Technology, Huazhong University of Science and Technology
No.1037 Luoyu Road, Hongshan District, Wuhan, Hubei 430074, China

Received:
December 5, 2025
Accepted:
April 2, 2026
Published:
July 20, 2026
Keywords:
scene text recognition, visual prompt learning, differential learning, textual prompts
Abstract

Scene text recognition (STR) in natural images remains highly challenging due to the large variations in character appearance across diverse real-world conditions, such as changes in font, color, layout, and background complexity—which hinder model generalization and remain insufficiently explored. To address this issue, we propose a visual prompt-guided differential learning (VPDL) framework designed to improve the generalization capability of STR models without requiring scene-specific fine-tuning. Inspired by the human ability to reference prior visual knowledge when recognizing text, VPDL introduces a set of character-level visual prompts that guide the model in perceiving appearance variations among characters. Built upon these prompts, we develop a local-to-global differential learning strategy that enhances patch-level representations and aligns global features with character cues while preserving scene-specific information. Additionally, to mitigate exposure bias in autoregressive decoding, we replace conventional label inputs with context-aware textual prompts, encouraging the decoder to better utilize textual cues embedded in image features. Extensive experiments on widely used benchmarks and real-world datasets demonstrate the effectiveness of VPDL.

Prompt-guided STR model

Prompt-guided STR model

Cite this article as:
Y. Niu, J. An, and S. Zou, “VPDL: Visual Prompt-Guided Differential Learning for Generalizable Scene Text Recognition,” J. Adv. Comput. Intell. Intell. Inform., Vol.30 No.4, pp. 1306-1317, 2026.
Data files:
References
  1. [1] Z. Zou, K. Chen, Z. Shi, Y. Guo, and J. Ye, “Object Detection in 20 Years: A Survey,” Proc. of the IEEE, Vol.111, Issue 3, pp. 257-276, 2023. https://doi.org/10.1109/JPROC.2023.3238524
  2. [2] Y. Chen, X. Yuan, J. Wang, R. Wu, X. Li, Q. Hou, and M.-M. Cheng, “YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.47, Issue 6, pp. 4240-4252, 2025. https://doi.org/10.1109/TPAMI.2025.3538473
  3. [3] M. Fang, X. Rui, H. Cheng, X. Liu, J. She, Y. Du, and H. Tan, “Small Object Detection Algorithm Based on Improved Attention Mechanism and Feature Fusion of YOLOv8,” J. Adv. Comput. Intell. Intell. Inform., Vol.29, No.4, pp. 941-955, 2025. https://doi.org/10.20965/jaciii.2025.p0941
  4. [4] J. Wan, S. Song, W. Yu, Y. Liu, W. Cheng, F. Huang, X. Bai, C. Yao, and Z. Yang, “OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 15641-15653, 2024. https://doi.org/10.1109/CVPR52733.2024.01481
  5. [5] B. Zhang, H. Xie, Z. Gao, and Y. Wang, “Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 28358-28368, 2024. https://doi.org/10.1109/CVPR52733.2024.02679
  6. [6] Q. Jiang, J. Wang, D. Peng, C. Liu, and L. Jin, “Revisiting Scene Text Recognition: A Data Perspective,” 2023 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 20486-20497, 2023. https://doi.org/10.1109/ICCV51070.2023.01878
  7. [7] C. Luo, C. Cheng, Q. Zheng, and C. Yao, “GeoLayoutLM: Geometric Pre-training for Visual Information Extraction,” 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 7092-7101, 2023. https://doi.org/10.1109/CVPR52729.2023.00685
  8. [8] R. Wang, Y. Xue, and L. Jin, “DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations,” Proc. of the AAAI Conf. on Artificial Intelligence, Vol.38, No.6, pp. 5563-5571, 2024. https://doi.org/10.1609/aaai.v38i6.28366
  9. [9] M. Peng, K. Chen, X. Guo, Q. Zhang, H. Zhong, M. Zhu, and H. Yang, “Diffusion Models for Intelligent Transportation Systems: A Survey,” IEEE Trans. on Intelligent Transportation Systems, Vol.26, Issue 12, pp. 21526-21543, 2025. https://doi.org/10.1109/TITS.2025.3613178
  10. [10] M. Bakirci, “Advanced aerial monitoring and vehicle classification for intelligent transportation systems with YOLOv8 variants,” J. of Network and Computer Applications, Vol.237, Article No.104134, 2025. https://doi.org/10.1016/j.jnca.2025.104134
  11. [11] M. A. Jan, M. Adil, B. Brik, S. Harous, and S. Abbas, “Making Sense of Big Data in Intelligent Transportation Systems: Current Trends, Challenges and Future Directions,” ACM Comput. Surv., Vol.57, Issue 8, Article No.197, 2025. https://doi.org/10.1145/3716371
  12. [12] C. Zhang, Y. Tao, K. Du, W. Ding, B. Wang, J. Liu, and W. Wang, “Character-Level Street View Text Spotting Based on Deep Multisegmentation Network for Smarter Autonomous Driving,” IEEE Trans. on Artificial Intelligence, Vol.3, Issue 2, pp. 297-308, 2022. https://doi.org/10.1109/TAI.2021.3116216
  13. [13] L. Wu, C. Zhang, J. Liu, J. Han, J. Liu, E. Ding, and X. Bai, “Editing Text in the Wild,” Proc. of the 27th ACM Int. Conf. on Multimedia, pp. 1500-1508, 2019. https://doi.org/10.1145/3343031.3350929
  14. [14] B. Shi, X. Bai, and C. Yao, “An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.39, Issue 11, pp. 2298-2304, 2017. https://doi.org/10.1109/TPAMI.2016.2646371
  15. [15] S. Fang, H. Xie, Y. Wang, Z. Mao, and Y. Zhang, “Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 7094-7103, 2021. https://doi.org/10.1109/CVPR46437.2021.00702
  16. [16] D. Bautista and R. Atienza, “Scene Text Recognition with Permuted Autoregressive Sequence Models,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 178-196, 2022. https://doi.org/10.1007/978-3-031-19815-1_11
  17. [17] Y. Du, Z. Chen, C. Jia, X. Yin, T. Zheng, C. Li, Y. Du, and Y.-G. Jiang, “SVTR: Scene Text Recognition with a Single Visual Model,” arXiv preprint, arXiv:2205.00159, 2022. https://doi.org/10.48550/arXiv.2205.00159
  18. [18] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” Proc. of the 23rd Int. Conf. on Machine Learning, pp. 369-376, 2006. https://doi.org/10.1145/1143844.1143891
  19. [19] Z. Cheng, F. Bai, Y. Xu, G. Zheng, S. Pu, and S. Zhou, “Focusing Attention: Towards Accurate Text Recognition in Natural Images,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 5086-5094, 2017. https://doi.org/10.1109/ICCV.2017.543
  20. [20] F. Xue, J. Sun, Y. Xue, Q. Wu, L. Zhu, X. Chang, and S.-C. Cheung, “Attention Guidance by Cross-Domain Supervision Signals for Scene Text Recognition,” IEEE Trans. on Image Processing, Vol.34, pp. 717-728, 2025. https://doi.org/10.1109/TIP.2024.3523799
  21. [21] J. Lee, S. Park, J. Baek, S. J. Oh, S. Kim, and H. Lee, “On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 2326-2335, 2020. https://doi.org/10.1109/CVPRW50498.2020.00281
  22. [22] J. Xu, Y. Wang, H. Xie, and Y. Zhang, “OTE: Exploring Accurate Scene Text Recognition Using One Token,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 28327-28336, 2024. https://doi.org/10.1109/CVPR52733.2024.02676
  23. [23] W. Hu, X. Cai, J. Hou, S. Yi, and Z. Lin, “GTC: Guided Training of CTC Towards Efficient and Accurate Scene Text Recognition,” Proc. of the AAAI Conf. on Artificial Intelligence, Vol.34, No.7, pp. 11005-11012, 2020. https://doi.org/10.1609/aaai.v34i07.6735
  24. [24] M. Yousef and T. E. Bishop, “OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by Learning to Unfold,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 14698-14707, 2020. https://doi.org/10.1109/CVPR42600.2020.01472
  25. [25] T. Wang, Y. Zhu, L. Jin, D. Peng, Z. Li, M. He, Y. Wang, and C. Luo, “Implicit Feature Alignment: Learn to Convert Text Recognizer to Text Spotter,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 5969-5978, 2021. https://doi.org/10.1109/CVPR46437.2021.00591
  26. [26] Y. Du, C. Li, R. Guo, X. Yin, W. Liu, J. Zhou, Y. Bai, Z. Yu, Y. Yang, Q. Dang, and H. Wang, “PP-OCR: A Practical Ultra Lightweight OCR System,” arXiv preprint, arXiv:2009.09941, 2020. https://doi.org/10.48550/arXiv.2009.09941
  27. [27] P. Wang, C. Da, and C. Yao, “Multi-granularity Prediction for Scene Text Recognition,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 339-355, 2022. https://doi.org/10.1007/978-3-031-19815-1_20
  28. [28] Y. Wang, H. Xie, S. Fang, J. Wang, S. Zhu, and Y. Zhang, “From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network,” 2021 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 14174-14183, 2021. https://doi.org/10.1109/ICCV48922.2021.01393
  29. [29] Z. Wan, M. He, H. Chen, X. Bai, and C. Yao, “TextScanner: Reading Characters in Order for Robust Scene Text Recognition,” Proc. of the AAAI Conf. on Artificial Intelligence, pp. 12120-12127, 2020. https://doi.org/10.1609/aaai.v34i07.6891
  30. [30] D. Yu, X. Li, C. Zhang, T. Liu, J. Han, J. Liu, and E. Ding, “Towards Accurate Scene Text Recognition with Semantic Reasoning Networks,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 12110-12119, 2020. https://doi.org/10.1109/CVPR42600.2020.01213
  31. [31] Z. Zhao, J. Tang, C. Lin, B. Wu, C. Huang, H. Liu, X. Tan, Z. Zhang, and Y. Xie, “Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 15567-15576, 2024. https://doi.org/10.1109/CVPR52733.2024.01474
  32. [32] B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “ASTER: An Attentional Scene Text Recognizer with Flexible Rectification,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.41, Issue 9, pp. 2035-2048, 2019. https://doi.org/10.1109/TPAMI.2018.2848939
  33. [33] Z. Cheng, F. Bai, Y. Xu, G. Zheng, S. Pu, and S. Zhou, “Focusing Attention: Towards Accurate Text Recognition in Natural Images,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 5086-5094, 2017. https://doi.org/10.1109/ICCV.2017.543
  34. [34] J. Baek, G. Kim, J. Lee, S. Park, D. Han, S. Yun, S. J. Oh, and H. Lee, “What is Wrong with Scene Text Recognition Model Comparisons? Dataset and Model Analysis,” 2019 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 4714-4722, 2019. https://doi.org/10.1109/ICCV.2019.00481
  35. [35] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Proc. of the Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol.1, pp. 4171-4186, 2019. https://doi.org/10.18653/v1/N19-1423
  36. [36] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” Proc. of the 34th Int. Conf. on Neural Information Processing Systems (NIPS’20), pp. 1877-1901, 2020. https://doi.org/10.18653/v1/N19-1423
  37. [37] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” OpenAI Blog, 2019.
  38. [38] A. Bar, Y. Gandelsman, T. Darrell, A. Globerson, and A. A. Efros, “Visual Prompting via Image Inpainting,” Advances in Neural Information Processing Systems, pp. 25005-25017, 2022.
  39. [39] J. Zhu and G. Pang, “Toward Generalist Anomaly Detection via In-Context Residual Learning with Few-Shot Sample Prompts,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 17826-17836, 2024. https://doi.org/10.1109/CVPR52733.2024.01688
  40. [40] F. Li, Q. Jiang, H. Zhang, T. Ren, S. Liu, X. Zou, H. Xu, H. Li, J. Yang, C. Li, L. Zhang, and J. Gao, “Visual in-Context Prompting,” 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 12861-12871, 2024. https://doi.org/10.1109/CVPR52733.2024.01222
  41. [41] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16times16 Words: Transformers for Image Recognition at Scale,” arXiv preprint, arXiv:2010.11929, 2020. https://doi.org/10.48550/arXiv.2010.11929
  42. [42] A. Mishra, K. Alahari, and C. Jawahar, “Scene Text Recognition using Higher Order Language Priors,” British Machine Vision Conf. (BMVC2), pp. 127.1-127.11, 2012. https://doi.org/10.5244/C.26.127
  43. [43] D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. Gomez i Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. A. Almazàn, and L.-P. de las Heras, “ICDAR 2013 Robust Reading Competition,” 2013 12th Int. Conf. on Document Analysis and Recognition, pp. 1484-1493, 2013. https://doi.org/10.1109/ICDAR.2013.221
  44. [44] D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny, “ICDAR 2015 Competition on Robust Reading,” 2015 13th Int. Conf. on Document Analysis and Recognition (ICDAR), pp. 1156-1160, 2015. https://doi.org/10.1109/ICDAR.2015.7333942
  45. [45] K. Wang, B. Babenko, and S. Belongie, “End-to-end scene text recognition,” 2011 IEEE Int. Conf. on Computer Vision, pp. 1457-1464, 2011. https://doi.org/10.1109/ICCV.2011.6126402
  46. [46] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing Text with Perspective Distortion in Natural Scenes,” 2013 IEEE Int. Conf. on Computer Vision, pp. 569-576, 2013. https://doi.org/10.1109/ICCV.2013.76
  47. [47] A. Risnumawan, P. Shivakumara, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” Expert Systems with Applications, Vol.41, Issue 18, pp. 8027-8048, 2014. https://doi.org/10.1016/j.eswa.2014.07.008
  48. [48] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, “Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 303-321, 2022. https://doi.org/10.1007/978-3-031-19815-1_18
  49. [49] A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie, “COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images,” arXiv preprint, arXiv:1601.07140, 2016. https://doi.org/10.48550/arXiv.1601.07140
  50. [50] C. K. Chng, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding et al., “ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text – RRC-ArT,” 2019 Int. Conf. on Document Analysis and Recognition (ICDAR), pp. 1571-1576, 2019. https://doi.org/10.1109/ICDAR.2019.00252
  51. [51] Y. Zhang, L. Gueguen, I. Zharkov, P. Zhang, K. Seifert, and B. Kadlec, “Uber-text: A large-scale dataset for optical character recognition from street-level imagery,” SUNw: Scene Understanding Workshop-CVPR, Vol.2017, 2017.
  52. [52] F. Sheng, Z. Chen, and B. Xu, “NRTR: A No-Recurrence Sequence-to-Sequence Model for Scene Text Recognition,” 2019 Int. Conf. on Document Analysis and Recognition, pp. 781-786, 2019. https://doi.org/10.1109/ICDAR.2019.00130
  53. [53] X. Yue, Z. Kuang, C. Lin, H. Sun, and W. Zhang, “RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition,” A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm (Eds.), “Computer Vision – ECCV 2020,” Lecture Notes in Computer Science, Vol.12364, pp. 135-151, 2020. https://doi.org/10.1007/978-3-030-58529-7_9
  54. [54] B. Na, Y. Kim, and S. Park, “Multi-modal Text Recognition Networks: Interactive Enhancements Between Visual and Semantic Features,” S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), “Computer Vision – ECCV 2022,” Lecture Notes in Computer Science, Vol.13688, pp. 446-463, 2022. https://doi.org/10.1007/978-3-031-19815-1_26
  55. [55] W. Xu, E. Eli, A. Aysa, X. Xu, and K. Ubul, “ASANet: Scene Text Recognition with Alternate Self-Attention,” 2025 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP), 2025. https://doi.org/10.1109/ICASSP49660.2025.10889330
  56. [56] C. Da, P. Wang, and C. Yao, “Multi-granularity prediction with learnable fusion for scene text recognition,” Int. J. of Computer Vision, Vol.134, No.1, Article No.47, 2026. https://doi.org/10.1007/s11263-025-02653-7
  57. [57] S. Zhao, R. Quan, L. Zhu, and Y. Yang, “CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-Trained Vision-Language Model,” IEEE Trans. on Image Processing, Vol.33, pp. 6893-6904, 2024. https://doi.org/10.1109/TIP.2024.3512354
  58. [58] R. Atienza, “Vision Transformer for Fast and Efficient Scene Text Recognition,” J. Lladós, D. Lopresti, and S. Uchida, “Document Analysis and Recognition – ICDAR 2021,” Lecture Notes in Computer Science, Vol.12821, pp. 319-334, 2021. https://doi.org/10.1007/978-3-030-86549-8_21
  59. [59] J. Baek, Y. Matsui, and K. Aizawa, “What If We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition with Fewer Labels,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 3112-3121, 2021. https://doi.org/10.1109/CVPR46437.2021.00313
  60. [60] M. Yang, M. Liao, P. Lu, J. Wang, S. Zhu, H. Luo, Q. Tian, and X. Bai, “Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text Recognition,” Proc. of the 30th ACM Int. Conf. on Multimedia (MM’22), pp. 4214-4223, 2022. https://doi.org/10.1145/3503161.3547784
  61. [61] X. Yang, Z. Qiao, J. Wei, D. Yang, and Y. Zhou, “Masked and Permuted Implicit Context Learning for Scene Text Recognition,” IEEE Signal Processing Letters, Vol.31, pp. 964-968, 2024. https://doi.org/10.1109/LSP.2024.3381893
  62. [62] Z. Wang, H. Xie, Y. Wang, J. Xu, B. Zhang, and Y. Zhang, “Symmetrical Linguistic Feature Distillation with CLIP for Scene Text Recognition,” Proc. of the 31st ACM Int. Conf. on Multimedia (MM’23), pp. 509-518, 2023. https://doi.org/10.1145/3581783.3611769

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Sep. 04, 2026