返回
Fully Convolutional Network for Multiscale Temporal Action Proposals
DOI:10.1109/TMM.2018.2839534.png)
摘要
En 中文
Similar to the function of object proposals in localizing objects within images, temporal action proposals can facilitate the extraction of semantic segments and simplify the computations required for temporal action localization in untrimmed videos. In this paper, we propose a fully convolutional network to identify multistale temporal action proposals (FCN-TAP) that utilizes only the temporal convolutions to retrieve accurate action proposals for video sequences. Using gated linear units, our network enables simple but powerful inferences, and by parallelizing the computations, it significantly improves performances compared with previous recurrent models. To capture more temporal contexts with fewer parameters, we apply dilated convolutions to expand the receptive fields of our network. Moreover, we divide the receptive fields into multiple scale ranges and then refine the corresponding temporal boundaries using duration regression at each scale. To generate suitable segments with arbitrary durations for training, we design a new strategy to select sampled candidates within the corresponding scale range. The power of our method is demonstrated through experiments on the THUMOS'14 and ActivityNet datasets, where FCN-TAP performs better and achieves a remarkable speedup compared to other state-of-the-art methods. Additional experiments show that our method generates high-quality proposals and improves the localization stage of existing action detection pipelines.
Keyword:
Temporal convolution
receptive field
multiple scale ranges
duration regression
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.7
论文数:
4.5K
被引数:
2.4W
机构
引用论文
The role of a critical left fronto-temporal network with its right-hemispheric homologue in syntactic learning based on word category information基于单词类别信息的关键左前-时间网络及其右半球同源物在句法学习中的作用

