arrow
Return

Interactive Image Caption Generation Reflecting User Intent from Trace Using a Diffusion Language Model

delete2025-11-01
delete0
PRE
AI
H
Hirano, Satoko *
I
Ichiro Kobayashi
DOI:10.20965/jaciii.2025.p1417delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This study proposes an image captioning method designed to incorporate user-specific explanatory intentions into the generated text, as signaled by the user's trace on the image. We extract areas of interest from dense sections of the trace, determine the order of explanations by tracking changes in the pen-tip coordinates, and assess the degree of interest in each area by analyzing the time spent on them. Additionally, a diffusion language model is utilized to generate sentences in a non-autoregressive manner, allowing control over sentence length based on the temporal data of the trace. In the actual caption generation task, the proposed method achieved higher string similarity than conventional methods, including autoregressive models, and successfully captured user intent from the trace and faithfully reflected it in the generated text.
Keywords:
deep learning
natural language processing
image captioning
diffusion model
interaction

Journal

J
Journal of Advanced Computational Intelligence and Intelligent Informatics
IF:
0.8
Papers:
87
Citations:
626

Organization

O
ochanomizu university
Scholars:
985
Papers: 848
Citations: 2