arrow
Return

Learning semantic consistency for audio-visual zero-shot learning

delete2025-04-17
delete0
delete
OA
AI
X
Xiaoyong Li
J
Jing Yang *
Y
Yuling Chen
W
Wei Zhang
X
Xiaoli Ruan
李澄江 (Chengjiang Li)
Z
Zhidong Su
DOI:10.1007/s10462-025-11228-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Audio-visual zero-shot learning requires an understanding of the relationship between audio and visual information to determine unseen classes. Despite many efforts and significant progress in the field, many existing methods tend to focus on learning strong representations, neglecting the semantic consistency between audio and video as well as the inherent hierarchical structure of the data. To address these issues, we propose Learning Semantic Consistency for Audio-Visual Zero-shot Learning. Specifically, we employ an attention mechanism to enhance cross-modal information interactions, aiming to capture the semantic consistency between audio and visual data. Meanwhile, we introduce a hyperbolic space to model the hierarchical structure of the data itself. Moreover, the proposed approach includes a novel loss function that considers the relationships between input modalities, reducing the distance between features of different modalities. To evaluate the proposed method, we test it on three benchmark datasets \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\hbox {VGGSound-GZS}{{\textrm{L}}<^>{cls}}$$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\hbox {UCF-GZS}{{\textrm{L}}<^>{cls}}$$\end{document}, and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\hbox {ActivityNet-GZS}{{\textrm{L}}<^>{cls}}$$\end{document}. Extensive experimental results show that the proposed method achieves state-of-the-art performance on all three datasets. For example, on the \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\hbox {UCF-GZS}{{\textrm{L}}<^>{cls}}$$\end{document} dataset, the harmonic mean is improved by 5.7%. Code and data available at https://github.com/ybyangjing/LSC-AVZSL.
Keywords:
Audio-visual zero-shot learning
Video classification
Attention mechanism
Hyperbolic geometry
Semantic consistency

Journal

Artificial Intelligence Review cover
Artificial Intelligence Review
IF:
13.9
Papers:
6.1K
Citations:
1.9W

Organization

O
Oklahoma State Univ
Scholars:
405
Papers: 301
Citations: 78