arrow
返回

Practical Video Object Detection via Feature Selection and Aggregation

delete2026-01-30
delete0
PRE
AI
Y
Yuheng Shi
T
Tong Zhang
X
Xiaojie Guo *
DOI:10.1007/s11263-025-02700-3delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Compared with still image object detection, video object detection (VOD) needs to particularly concern the high across-frame variation in object appearance, and the diverse deterioration in some frames. In principle, the detection in a certain frame of a video can benefit from information in other frames. Thus, how to effectively aggregate features across different frames is key to the target problem. Most of contemporary aggregation methods are tailored for two-stage detectors, suffering from high computational costs due to the dual-stage nature. On the other hand, although one-stage detectors have made continuous progress in handling static images, their applicability to VOD lacks sufficient exploration. To tackle the above issues, this study invents a very simple yet potent strategy of feature selection and aggregation, gaining significant accuracy at marginal computational expense. Concretely, for cutting the massive computation and memory consumption from the dense prediction characteristic of one-stage object detectors, we first condense candidate features from dense prediction maps. Then, the relationship between a target frame and its reference frames is evaluated to guide the aggregation. Comprehensive experiments and ablation studies are conducted to validate the efficacy of our design, and showcase its advantage over other cutting-edge VOD methods in both effectiveness and efficiency. Notably, our model reaches a new record performance, i.e., 93.0% AP50 at over 30 FPS on the ImageNet VID dataset on a single 3090 GPU, making it a compelling option for large-scale or real-time applications. The implementation is simple, and accessible at https://github.com/YuHengsss/YOLOV .
Keyword:
Object detection
Video object detection
Feature selection
Feature aggregation

期刊

International Journal of Computer Vision 封面图
International Journal of Computer Vision
IF:
9.3
论文数:
3.9K
被引数:
2.8W

机构

C
College of Intelligence and Computing
学者数:
152
论文数: 69
被引数: 1
引用论文

引用论文

A ConvNet for the 2020s21世纪20年代的ConvNet
err2022-06-01
err0
PREAI
errZhuang Liu; Hanzi Mao; Chao-Yuan Wu; Christoph Feichtenhofer; Trevor Darrell; Saining Xie
err分享
err收藏
TIDE: A General Toolbox for Identifying Object Detection Errors
err2020-12-03
err0
PREAI
errDaniel Bolya; Sean Foley; James Hays; Judy Hoffman
err分享
err收藏
Deep Feature Flow for Video Recognition
err2017-07-01
err0
errOAAI
errXizhou Zhu; Yuwen Xiong; Jifeng Dai; Lu Yuan; Yichen Wei
err分享
err收藏
End-to-End Object Detection with Transformers使用Transformers进行端到端对象检测
err2020-11-03
err0
PREAI
errNicolas Carion; Francisco Massa; Gabriel Synnaeve; Nicolas Usunier; Alexander Kirillov; Sergey Zagoruyko
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
学者 查看更多内容