1
Return

RD-VTA: Rule-Data Guided Video-to-Audio Generation for Fine-Grained Footstep Sound

delete2026-06-12
delete0
PRE
AI
Q
Qiutang Qi
H
Haonan Cheng
H
Hengyan Huang
叶龙 cover
叶龙 (Long Ye)
S
Shaobin Li
DOI:10.1109/tmm.2026.3703315delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
It is challenging to implement visually guided fine-grained footstep sounds based on a limited number of samples in complex scenes. This is due to the interference of redundant information in the complex background of the visual scene for audiovisual mapping. As well as the complex coupling of sound features makes audiovisual fine-grained linear mapping difficult. To address the mentioned problems, we propose an automated video-to-audio generation method (RD-VTA) for footstep sound that incorporates data-driven and rule-based modelling approaches. First, we design a data-driven masked footstep sound generation network (DM-AFSG) to acquire audiovisual temporal concordance. The network is capable of separating visual sound objects, reducing background redundant interference, and generating initial target sounds that capture temporal cues. Secondly, a rule-based fine-grained footstep sound adjustment method (RT-AFSG) is designed based on visual guides such as material, motion type and displacement distance. The proposed RT-AFSG effectively achieves diverse sounds with a limited number of sound samples through sound texture analysis and modification. Moreover, it constructs the mapping relationship between different visual cues and footstep sounds, and realizes the fine variation of footstep sounds. To adequately validate the effectiveness of the method in terms of audiovisual temporal consistency and content granularity, we perform objective synchronization metrics and subjective human evaluation on the footsteps audiovisual dataset VAFoot. The experimental results show that the method obtains an average of 5% improvement in sound synchronization performance and significantly outperforms several existing methods in terms of sound content granularity.
Keywords:
Data-driven method
fine-grained footstep sound
rule-based modeling
video-to-audio generation

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.4K
Citations:
2.4W

Organization

C
communication university of china
Scholars:
404
Papers: 211
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers