返回
Knowledge distilled pre-training model for vision-language-navigation
DOI:10.1007/s10489-022-03779-8.png)
摘要
En 中文
Vision-language-navigation(VLN) is a challenging task that requires a robot to autonomously move to a destination based on visual observation following a human's natural language instructions. To improve the performance and generalization ability, the pre-training model based on the transformer is used instead of the traditional methods. However, the pre-training model is not suitable for sustainable computing and practical application because of its complex computations and large amount of hardware occupation. Therefore, we propose a slight pre-training model through knowledge distillation. Through knowledge distillation, the plenty of knowledge encoded in a large teacher model can be well transferred to a small student model, which greatly reduces the model parameters and inference time while maintaining the original performance. In the experiments, the model size is reduced by 87%, and the average inference time is reduced by approximately 86%. It can be trained and run much faster. At the same time, 95% performance of the original model was maintained, which is still better than the traditional VLN models.
Keyword:
Natural language processing
Computer vision
Cross-modality
Deep learning
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
Multi-teacher knowledge distillation for compressed video action recognition based on deep learning基于深度学习的压缩视频动作识别多教师知识提炼

