arrow
Return

Dual-path temporal map optimization for make-up temporal video grounding

delete2024-05-03
delete2
PRE
AI
J
Jiaxiu Li
K
Kun Li
李家 (Jia Li)
G
Guoliang Chen
王萌 (Meng Wang)
D
Dan Guo *
DOI:10.1007/s00530-024-01340-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Make-up temporal video grounding (MTVG) aims to localize the target video segment, which is semantically related to a sentence describing a make-up activity in a make-up video. Compared with the general video grounding, MTVG focuses on meticulous actions and changes on the face. The make-up instruction step, usually involving detailed differences in products and facial areas, is more fine-grained than general activities (e.g., cooking activity and furniture assembly). Thus, existing general approaches may not effectively locate the target activity effectually due to the lack of fine-grained semantic cues for the make-up semantic comprehension. To tackle this issue, we propose an effective proposal-based framework named Dual-Path Temporal Map Optimization Network to capture fine-grained multimodal semantic details of make-up activities. We extract both query-agnostic and query-guided features to construct two proposal sets and use specific evaluation methods for the two sets. Different from the commonly used single structure in previous methods, our dual-path structure can mine more semantic information in make-up videos and distinguish fine-grained actions well. These two candidate sets represent the cross-modal makeup video-text similarity and multi-modal fusion relationship, complementing each other. Therefore, the joint prediction of these sets will enhance the accuracy of video timestamp prediction. Comprehensive experiments on the YouMakeup dataset demonstrate our proposed dual structure excels in fine-grained semantic comprehension. The source code will be available at: https://github.com/lijiaxiuHFUT/DPTMO.
Keywords:
Video understanding
Make-up temporal video grounding
Proposal generation
2D temporal map

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.7K
Citations:
2.7K

Organization

H
hefei university of technology
Scholars:
2.5W
Papers: 1.7W
Citations: 35