Return
Emocoop: dynamic vision–language coupling for multi-label emotion classification in Dongba paintings
R
W
P
DOI:10.1007/s00371-026-04568-x.png)
Abstract
En 中文
Dongba paintings, a unique Naxi cultural heritage, are characterized by complex symbolism and diverse visual aesthetics. However, contemporary vision–language models exhibit insufficient domain generalization capabilities in few-shot emotion classification on these artworks, substantially impeding accurate multi-label emotion recognition. To address this critical challenge, we propose EmoCoOp, a dynamic vision–language coupling framework. It incorporates latent priors from visual semantics to guide prompt learning and employs a meta-network for adaptive visual prompt generation. A vision-language coupling mechanism facilitates deep multimodal integration, while a dual-path chromatic affective inference module models coloration and emotional expressions. Experimental results demonstrate EmoCoOp’s superior performance, achieving 75.73% mAP and 82.21% Recall@2, outperforming the second-ranked model by significant margins. This framework advances multi-label emotion classification for ethnic artworks and cross-modal understanding. The code is available at https://github.com/yang-easy/EmoCoOp .
Keywords:
Dongba paintings
Multi-label emotion classification
Visual-language coupling
Dynamic visual prompt generation
Journal
IF:
2.9
Papers:
4.5K
Citations:
6.5K
