Return
Knowledge-Driven Compositional Action Recognition
DOI:10.1016/j.patcog.2025.111452.png)
Abstract
En 中文
Human action often involves interaction with objects, so in action recognition, action labels can be defined by compositions of verbs and nouns. It is almost infeasible to collect and annotate enough training data for every possible composition in the real world. Therefore, the main challenge in compositional action recognition is to enable the model to understand action-objectscompositions that have not been seen during training. We propose a Knowledge-Driven Composition Modulation Model (KCMM), which constructs unseen action-objectscompositions to improve action recognition generalization. We first design a Grammar Knowledge-Driven Composition (GKC) module, which extracts the labels of verbs and nouns and their corresponding feature representations from compositional actions, and then modulates them under the guidance of grammatical rules to construct new action-objectsactions. Subsequently, to verify the rationality of the new action-objectsactions, we design a Common Knowledge-Driven Verification (CKV) module. This module extracts motion commonsense from ConceptNet and infuses it into the compositional labels to improve the comprehensiveness of the verification. It should be noted that GKC does not construct new videos, but directly composes verbs and nouns at the label and feature space to obtain new compositional action label-feature pairs. We conduct extensive experiments on Something-Else and NEU-I datasets, and our method significantly outperforms current state-of-the-art methods in both compositional settings and few-shot settings. The source code is available at https://github.com/XDLiuyyy/KCMM.
Keywords:
Compositional action recognition
Compositional learning
Knowledge-Driven
Few-shot
Knowledge graph
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W
Organization
No organization information available

