Return
An fMRI visual neural encoding method with multimodal large language model
DOI:10.1016/j.knosys.2025.114049.png)
Abstract
En 中文
• In summary, our contributions are primarily threefold:. • To our knowledge, we establish the first multimodal framework combining MLLM with fMRI visual neural encoding, introducing a systematic three-phase training paradigm specifically optimized for neural encoding tasks. • Building upon the Vicuna architecture, we develop an 8-billion-parameter foundation model that demonstrates dual advantages in parameter efficiency and task performance. Specifically, our method achieves 5th place in Algonauts 2023, establishing a new benchmark in large-scale visual encoding processing. • Ablation studies demonstrate that introducing the MLLM module yields a 2.87 % performance gain. Through our multi-stage training paradigm, we achieve a 2.61 % performance improvement by fine-tuning only 1.33 % of the parameters (Q-former) via parameter-efficient fine-tuning, achieving a balance between computational efficiency and performance enhancement.
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
Organization
No organization information available

