Return
A hierarchical active inference framework for stable robotic control
DOI:10.1016/j.eswa.2025.130903.png)
Abstract
En 中文
This paper presents Active Inference-Visuomotor Policy Learning (AIF-VPL), a novel active inference framework that bridges neuroscience principles with robotic imitation learning. While current approaches struggle to balance stability and adaptability, our architecture demonstrates how cortical-cerebellar-spinal computational principles can resolve this challenge. At the cortical level, a hybrid Conv-xLSTM network integrated with a Multimodal Attention Module (MAM) processes spatiotemporal visual-proprioceptive inputs for task planning. The cerebellar-inspired mid-level employs precision-weighted variational autoencoders to implement active inference, which minimizes sensory prediction errors through iterative action refinement, improving stability (35% jerk reduction). Finally, a spinal-level xLSTM-Transformer network ensures low-latency motor execution through structured action sequences. Evaluated on three manipulation tasks, Drag, Transfer, and Push-T, AIFVPL achieves 93-100% success rates, outperforming diffusion policies and behavioral cloning baselines. Ablation studies confirm that each neurobiologically inspired component is essential: the MAM yields temporally coherent representations, while active inference reduces trajectory jerk by 35%. This work presents the first deployable implementation of hierarchical active inference in robotics, offering a principled framework that bridges neurorobotics and computational neuroscience.
Keywords:
Hierarchical control
Imitation learning
Multimodal fusion
Active inference
Robotic sensorimotor tasks
Visuomotor policy learning
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

