arrow
Return

Disentangled implicit and explicit multimodal intent learning for sequential recommendation

delete2026-08-18
delete0
PRE
AI
K
Kang Yang
蔡德胜 cover
蔡德胜 (Desheng Cai) *
Y
Ying He *
DOI:10.1007/s00530-026-02591-5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Sequential recommendation aims to predict the next item based on users’ historical interactions. While recent studies have highlighted the importance of modeling user intents, most existing methods either infer intents implicitly from interaction data or rely on large language models (LLMs) to extract intents from textual information, overlooking the rich multimodal signals available in real-world scenarios. Although multimodal data provides complementary information for intent discovery, directly utilizing it introduces challenges such as semantic entanglement across modalities and the lack of effective explicit intent extraction mechanisms. In this paper, we propose DMIRec, a disentangled multimodal intent learning framework for sequential recommendation. DMIRec jointly leverages explicit intents extracted by a multimodal large language model (MLLM) and implicit intent prototypes learned via vector quantization. Specifically, we employ an MLLM to perform explicit cross-modal intent extraction, generating semantically meaningful intent priors. To model fine-grained intent structures, we design a disentangled implicit intent learning module that separates modality-shared and modality-specific representations via alignment and orthogonality constraints. Furthermore, a vector quantization mechanism is introduced to learn compact and representative intent prototypes. Finally, we develop an intent-aware attention mechanism to integrate both explicit and implicit intents for enhanced user representation learning. Extensive experiments on three real-world datasets demonstrate that DMIRec consistently outperforms state-of-the-art methods, validating the effectiveness of combining explicit and implicit intent modeling with multimodal disentanglement.
Keywords:
Sequential recommendation
User intent modeling
Representation disentanglement
Vector quantization
Multimodal large language models

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.7K
Citations:
2.7K

Organization

S
School of Computer Science and Engineering
Scholars:
1.3K
Papers: 586
Citations: 2