arrow
返回

Multiple instance deep learning for weakly-supervised visual object tracking

delete2020-05-01
delete3
PRE
AI
K
Kaining Huang
Y
Yan Shi *
F
Fuqi Zhao
Z
Zijun Zhang
S
Shanshan Tu
DOI:10.1016/j.image.2020.115807delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Intelligently tracking objects with varied shapes, color, lighting conditions, and backgrounds is an extremely useful application in many HCI applications, such as human body motion capture, hand gesture recognition, and virtual reality (VR) games. However, accurately tracking different objects under uncontrolled environments is a tough challenge due to the possibly dynamic object parts, varied lighting conditions, and sophisticated backgrounds. In this work, we propose a novel semantically-aware object tracking framework, wherein the key is weakly-supervised learning paradigm that optimally transfers the video-level semantic tags into various regions. More specifically, give a set of training video clips, each of which is associated with multiple video-level semantic tags, we first propose a weakly-supervised learning algorithm to transfer the semantic tags into various video regions. The key is a MIL (Zhong et al., 2020) [1]-based manifold embedding algorithm that maps the entire video regions into a semantic space, wherein the video-level semantic tags are well encoded. Afterward, for each video region, we use the semantic feature combined with the appearance feature as its representation. We designed a multi-view learning algorithm to optimally fuse the above two types of features. Based on the fused feature, we learn a probabilistic Gaussian mixture model to predict the target probability of each candidate window, where the window with the maximal probability is output as the tracking result. Comprehensive comparative results on a challenging pedestrian tracking task as well as the human hand gesture recognition have demonstrated the effectiveness of our method. Moreover, visualized tracking results have shown that non-rigid objects with moderate occlusions can be well localized by our method.
Keyword:
Multiple instance learning (MIL)
Weakly-supervised
Object tracking
Multi-view feature learning
Gaussian mixture model
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

S
Signal Processing and Image Communication
IF:
2.7
论文数:
2.8K
被引数:
4.2K

机构

B
Bengbu University
学者数:
589
论文数: 343
被引数: 181
B
Beijing University of Technology
学者数:
2.8W
论文数: 2.1W
被引数: 2.7W
引用论文

引用论文

Multi-view kernel machine on single-view data
err2009-06-01
err25
PREAI
errWang, Zhe; Chen, Songcan
err分享
err收藏
Improved seam carving for video retargeting
err2008-08-01
err602
PREAI
errRubinstein, Michael; Shamir, Ariel; Avidan, Shai
err分享
err收藏
Phase‐Contact Engineering in Mono‐ and Bimetallic Cu‐Ni Co‐catalysts for Hydrogen Photocatalytic Materials
err2018-01-11
err0
PREAI
errMario J. Muñoz‐Batista; Debora Motta Meira; Gerardo Colón; Anna Kubacka; Marcos Fernández‐García
err分享
err收藏
Inflammatory bowel diseases activity in patients undergoing pelvic radiation therapy
err2017-02-01
err0
errOAAI
errPierre Annede; Thomas Seisen; Caroline Klotz; Renaud Mazeron; Pierre Maroun; Claire Petit; Eric Deutsch; Alberto Bossi; Christine Haie-Meder; Cyrus Chargari; Pierre Blanchard
err分享
err收藏
学者 查看更多内容