Return
Unlocking human intent perception through multimodal large models
DOI:10.1016/j.patcog.2025.112582.png)
Abstract
En 中文
• IntentMLM replaces text decoder with a linear layer for direct intent scoring. • Retrieval-augmented strategy boosts accuracy via multimodal sample fusion. • Text description models enrich modalities to enhance intent prediction robustness.
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

