arrow
Return

DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks

delete2026-01-01
delete0
PRE
AI
Y
Yinqi Li
常虹 cover
常虹 (Hong Chang)
R
Ruibing Hou
S
Shiguang Shan
陈熙霖 (Xilin Chen)
DOI:10.1109/TMM.2025.3623508delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Diffusion models have shown remarkable progress in various generative tasks such as image and video generation. This paper studies the problem of leveraging pretrained diffusion models for performing discriminative tasks. Specifically, we extend the discriminative capability of pretrained frozen generative diffusion models from the classification task (Li et al., 2023), (Clark et al., 2023) to the more complex object detection task, by “inverting” a pretrained layout-to-image diffusion model. To this end, a gradient-based discrete optimization approach for replacing the heavy prediction enumeration process, and a prior distribution model for making more accurate use of the Bayes' rule, are proposed respectively. Empirical results show that this method is on par with basic discriminative object detection baselines on COCO dataset. In addition, our method can greatly speed up the previous diffusion-based method (Li et al., 2023), (Clark et al., 2023) for classification without sacrificing accuracy. Code and models are available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/LiYinqi/DIVE</uri>.
Keywords:
Diffusion model
generative modeling
discriminative Task
object detection
visual recognition

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.4K
Citations:
2.4W

Organization

C
Chinese Academy of Sciences
Scholars:
3.9W
Papers: 1.5W
Citations: 58.4W