Return
ToM: Boosting TextVQA by capturing text-oriented keypoints
DOI:10.1016/j.knosys.2026.115480.png)
Abstract
En 中文
Text-based Visual Question Answering (TextVQA) requires recognizing and understanding the text in images, making text crucial for answer reasoning. Existing methods rely on external OCR systems that extract text indiscriminately, causing a disconnection between text extraction and reasoning and preventing correction of recognition errors during inference. In this paper, we propose a novel text-oriented unified framework that enables the model to dynamically refine and leverage question-relevant text throughout the reasoning process. It couples a trainable question-aware scene text spotter with an iteratively optimized answer generator under a unified optimization objective. The text spotter performs global text extraction while emphasizing question-relevant regions, and can be continuously refined via feedback from the reasoning stage, producing more task-oriented text than off-the-shelf OCR systems. Furthermore, to mitigate the influence of distracting text, our answer generator adopts a delivery-feedback mechanism that iteratively evaluates candidate answers to suppress irrelevant information. Extensive experiments on TextVQA and ST-VQA demonstrate the framework’s effectiveness, and validation on TextVideoQA shows it generalizes well to video scenarios.
Keywords:
TextVQA
scene text spotting
answer generation
unified framework
iterative refinement
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
Organization
No organization information available

