arrow
返回

Consistencydet: few-step self-consistent box denoising for efficient object detection

delete2026-08-07
delete0
PRE
AI
L
Lifan Jiang
王智慧 封面图
王智慧 (Zhihui Wang) *
C
Changmiao Wang
M
Ming Li
J
Jiaxu Leng
DOI:10.1007/s00371-026-04674-wdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
迭代框去噪 recently reformulated object detection as progressive noise-to-box refinement. However, diffusion-style detectors such as DiffusionDet inherit a key diffusion bottleneck: accurate inference often needs long reverse chains and repeated decoder calls. This paper proposes ConsistencyDet, a few-step self-consistent box denoising framework for closed-set object detection. Our motivation is to transfer the few-step sampling advantage of consistency models from generative modelling to discriminative detection. During training, Gaussian-corrupted ground-truth boxes and RoI features are decoded by a shared noise-conditioned detector. We introduce a detection-specific consistency objective along a probability-flow ordinary differential equation: adjacent noisy box states are decoded with shared weights and supervised by the same matched clean detection targets. At inference, random box proposals are refined in a few denoising steps, while Box-renewal replaces low-confidence proposals to keep the proposal distribution close to the training corruption process. On MS-COCO with ResNet-50, ConsistencyDet obtains 46.80 AP at two sampling steps and 6.892 FPS, exceeding DiffusionDet’s 46.54 AP at twenty steps while being about 7.2 times faster under the same RTX3080 protocol. Experiments on MS-COCO and LVIS with convolutional and transformer backbones show competitive accuracy with much lower long-chain sampling cost. Code, configurations, and reproduction scripts are released at https://github.com/Tankowa/ConsistencyDet .
Keyword:
Efficient object detection
Box denoising
Consistency models
Diffusion-style detection
Few-step sampling
Probability-flow ODE
Reproducible visual computing benchmark

期刊

Visual Computer 封面图
Visual Computer
IF:
2.9
论文数:
4.6K
被引数:
6.5K

机构

S
Shenzhen Research Institute of Big Data
学者数:
256
论文数: 352
被引数: 357
引用论文

引用论文

Object Detection in 20 Years: A Survey20年来的目标检测: 一项调查
err2023-03-01
err812
errOAAI
errZou, Zhengxia; Chen, Keyan; Shi, Zhenwei; Guo, Yuhong; Ye, Jieping
err分享
err收藏
The Pascal Visual Object Classes (VOC) ChallengePascal视觉对象课程 (VOC) 挑战
err2009-09-09
err9.0K
PREAI
errEveringham, Mark; Van Gool, Luc; Williams, Christopher K. I.; Winn, John; Zisserman, Andrew
err分享
err收藏
Diffusion Models in Vision: A Survey视觉中的扩散模型: 综述
err2023-09-01
err0
errOAAI
errFlorinel-Alin Croitoru; Vlad Hondru; Radu Tudor Ionescu; Mubarak Shah
err分享
err收藏
没有更多内容