arrow
Return

Underlying Semantic Diffusion for Effective and Efficient In-Context Learning

delete2026-07-02
delete0
PRE
AI
冀中 cover
冀中 (Zhong Ji)
W
Weilong Cao
Y
Yan Zhang
Y
Yanwei Pang
韩军功 (Jungong Han)
DOI:10.1109/TIP.2026.3707788delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Diffusion models have emerged as a powerful framework for tasks like image controllable generation and dense prediction. However, existing models often struggle to capture underlying semantics (e.g., edges, textures, shapes) and effectively utilize in-context learning, limiting their contextual understanding and image generation quality. Furthermore, high computational costs and slow inference speeds hinder their real-time applications. To address these challenges, we propose Underlying Semantic Diffusion (US-Diffusion), an enhanced diffusion model that improves underlying semantics learning, computational efficiency, and in-context learning capabilities on multi-task scenarios. We introduce Separate & Gather Adapter (SGA), which decouples input conditions for different tasks while sharing the architecture, enabling better in-context learning and generalization across diverse visual domains. We also present a Feedback-Aided Learning (FAL) framework, which leverages feedback signals to guide the model in capturing semantic details and dynamically adapting to task-specific contextual cues. Furthermore, we propose a plug-and-play Efficient Sampling Strategy (ESS) for dense sampling at time steps with high-noise levels, which aims at optimizing training and inference efficiency while maintaining strong in-context learning performance. Experimental results demonstrate that US-Diffusion outperforms the state-of-the-art method, achieving an average reduction of 7.47 in FID on Map2Image tasks and an average reduction of 0.026 in RMSE on Image2Map tasks, while achieving approximately <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$9.45\times $ </tex-math></inline-formula> faster inference speed. Our method also demonstrates superior training efficiency and in-context learning capabilities, excelling in new datasets and tasks, highlighting its robustness and adaptability across diverse visual domains. The source code will be released at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/dragon-cao/US-Diffusion</uri>
Keywords:
Diffusion model
in-context learning
underlying semantics
efficient sampling

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

T
tsinghua university
Scholars:
11.6W
Papers: 9.9W
Citations: 137
T
tianjin university
Scholars:
7.8W
Papers: 5.7W
Citations: 88