返回
Robust Visual Tracking via Convolutional Networks Without Training
DOI:10.1109/TIP.2016.2531283.png)
摘要
En 中文
Deep networks have been successfully applied to visual tracking by learning a generic representation offline from numerous training images. However, the offline training is time-consuming and the learned generic representation may be less discriminative for tracking specific objects. In this paper, we present that, even without offline training with a large amount of auxiliary data, simple two-layer convolutional networks can be powerful enough to learn robust representations for visual tracking. In the first frame, we extract a set of normalized patches from the target region as fixed filters, which integrate a series of adaptive contextual filters surrounding the target to define a set of feature maps in the subsequent frames. These maps measure similarities between each filter and useful local intensity patterns across the target, thereby encoding its local structural information. Furthermore, all the maps together form a global representation, via which the inner geometric layout of the target is also preserved. A simple soft shrinkage method that suppresses noisy values below an adaptive threshold is employed to de-noise the global representation. Our convolutional networks have a lightweight structure and perform favorably against several state-of- the-art methods on the recent tracking benchmark data set with 50 challenging videos.
Keyword:
Visual tracking
convolutional networks
deep learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
13.7
论文数:
1.0W
被引数:
8.4W
机构
引用论文
Detecting dead regions using psychophysical tuning curves: A comparison of simultaneous and forward masking使用心理物理调谐曲线检测死区: 同时掩蔽和正向掩蔽的比较

