arrow
Return

Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models

delete2026-01-01
delete0
PRE
AI
M
Michael Jungo *
A
Andreas Fischer
DOI:10.1007/978-3-032-09368-4_18delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Rule-based reinforcement learning has been gaining popularity ever since DeepSeek-R1 has demonstrated its success through simple verifiable rewards. In the domain of document analysis, reinforcement learning is not as prevalent, even though many downstream tasks may benefit from the emerging properties of reinforcement learning, particularly the enhanced reason capabilities. We study the effects of rule-based reinforcement learning with the task of Document Image Classification which is one of the most commonly studied downstream tasks in document analysis. We find that reinforcement learning tends to have better generalisation capabilities to out-of-distritbution data, which we examine in three different scenarios, namely out-of-distribution images, unseen classes and different modalities. Our code is available at https://github.com/jungomi/vision-finetune.
Keywords:
Document Image Classification
Vision Language Models
Large Language Models
Reinforcement Learning

Journal

D
DOCUMENT ANALYSIS AND RECOGNITION - ICDAR 2025 WORKSHOPS, PT I
IF:
0
Papers:
22
Citations:
0

Organization

U
university of fribourg
Scholars:
797
Papers: 376
Citations: 0