single-jc.php

JACIII Vol.30 No.4 pp. 1306-1317
(2026)

Research Paper:

VPDL: Visual Prompt-Guided Differential Learning for Generalizable Scene Text Recognition

Received:
December 5, 2025
Accepted:
April 2, 2026
Published:
July 20, 2026
Abstract

Scene text recognition (STR) in natural images remains highly challenging due to the large variations in character appearance across diverse real-world conditions, such as changes in font, color, layout, and background complexity—which hinder model generalization and remain insufficiently explored. To address this issue, we propose a visual prompt-guided differential learning (VPDL) framework designed to improve the generalization capability of STR models without requiring scene-specific fine-tuning. Inspired by the human ability to reference prior visual knowledge when recognizing text, VPDL introduces a set of character-level visual prompts that guide the model in perceiving appearance variations among characters. Built upon these prompts, we develop a local-to-global differential learning strategy that enhances patch-level representations and aligns global features with character cues while preserving scene-specific information. Additionally, to mitigate exposure bias in autoregressive decoding, we replace conventional label inputs with context-aware textual prompts, encouraging the decoder to better utilize textual cues embedded in image features. Extensive experiments on widely used benchmarks and real-world datasets demonstrate the effectiveness of VPDL.

Prompt-guided STR model

Prompt-guided STR model

Cite this article as:
, “VPDL: Visual Prompt-Guided Differential Learning for Generalizable Scene Text Recognition,” J. Adv. Comput. Intell. Intell. Inform., Vol.30 No.4, pp. 1306-1317, 2026.
Data files:

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Jul. 19, 2026