arrow
Return

Audio-guided self-supervised learning for disentangled visual speech representations

delete2024-06-25
delete0
PRE
AI
D
Dalu Feng
S
Shuang Yang *
S
Shiguang Shan
陈熙霖 (Xilin Chen)
DOI:10.1007/s11704-024-3787-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
4 ConclusionIn this paper, we propose a novel two-branch framework to learn the disentangled visual speech representations based on two particular observations. Its main idea is to introduce the audio signal to guide the learning of speech-relevant cues and introduce a bottleneck to restrict the speech-irrelevant branch from learning high-frequency and fine-grained speech cues. Experiments on both the word-level and sentence-level audio-visual speech datasets LRW and LRS2-BBC show the effectiveness. Our future work is to explore more explicit auxiliary tasks and constraints beyond the reconstruction task of the speech-relevant and irrelevant branch to improve further its ability of capturing speech cues in the video. Meanwhile, it's also a nice try to combine multiple types of knowledge representations [10] to further boost the obtained speech epresentations, which is also left for the future work.

Journal

Frontiers of Computer Science cover
Frontiers of Computer Science
IF:
4.6
Papers:
1.6K
Citations:
2.8K

Organization

C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704