arrow
Return

An fMRI visual neural encoding method with multimodal large language model

delete2025-06-27
delete0
PRE
AI
S
Shuxiao Ma
L
Linyuan Wang
L
Libin Hou
S
Senbao Hou
B
Bin Yan
DOI:10.1016/j.knosys.2025.114049delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• In summary, our contributions are primarily threefold:. • To our knowledge, we establish the first multimodal framework combining MLLM with fMRI visual neural encoding, introducing a systematic three-phase training paradigm specifically optimized for neural encoding tasks. • Building upon the Vicuna architecture, we develop an 8-billion-parameter foundation model that demonstrates dual advantages in parameter efficiency and task performance. Specifically, our method achieves 5th place in Algonauts 2023, establishing a new benchmark in large-scale visual encoding processing. • Ablation studies demonstrate that introducing the MLLM module yields a 2.87 % performance gain. Through our multi-stage training paradigm, we achieve a 2.61 % performance improvement by fine-tuning only 1.33 % of the parameters (Q-former) via parameter-efficient fine-tuning, achieving a balance between computational efficiency and performance enhancement.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

No organization information available