arrow
Return

Joint Underwater Image Enhancement and Captioning Through Multisupervised Chained Task Learning

delete2026-07-07
delete0
PRE
AI
H
Huanyu Li
L
Li Li
H
Hao Wang
W
Weibo Zhang
刘婧宇(LiuJingyu) (Jingyu Liu)
P
Peng Ren
DOI:10.1109/tgrs.2026.3710783delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Underwater vision expands the realm of computer vision from terrestrial to aquatic environments. To explore light-degraded and poorly comprehensible aquatic scenes, underwater image enhancement and underwater image captioning stand as two highly demanded tasks. Despite their potential to mutually reinforce each other, these two tasks have remained largely isolated in existing research. To bridge this gap, we formulate underwater image enhancement and underwater image captioning as a chained low-level-to-high-level vision–language learning problem, aiming to connect low-level visual restoration with high-level semantic generation within a unified framework. Different from conventional enhance-then-caption pipelines, the proposed formulation enables the enhancement process to be constrained not only by perceptual restoration objectives, but also by downstream semantic captioning supervision. To instantiate this formulation, we develop a task-coupled enhancement-captioning framework consisting of a color correction-guided dual-domain enhancement branch and a spatial cluster-aware captioning branch. We further develop a multisupervised chained task learning strategy that coordinates reconstruction supervision, differentiable quality-oriented proxy supervision, and CIDEr-based sequence-level reward supervision within a unified optimization process. In addition, we construct the underwater image enhancement and captioning (UIEC) dataset to provide the data basis and evaluation benchmark for joint UIEC. Evaluations on the UIEC benchmark demonstrate that the proposed joint framework achieves superior performance in both underwater image enhancement and underwater image captioning tasks, validating the effectiveness of the chained learning paradigm for collaborative underwater visual restoration and semantic generation.
Keywords:
Multisupervised chained task learning
underwater image captioning
underwater image enhancement

Journal

IEEE Transactions on Geoscience and Remote Sensing cover
IEEE Transactions on Geoscience and Remote Sensing
IF:
8.6
Papers:
2.1W
Citations:
10.7W

Organization

L
Laoshan Laboratory
Scholars:
2.7K
Papers: 2.1K
Citations: 1.3K
C
china university of petroleum (east china)
Scholars:
4.9K
Papers: 1.3K
Citations: 0