arrow
Return

VIMAR: vision-language informed malware analysis and reasoning model

delete2026-03-27
delete0
delete
OA
AI
S
Shiting Xu *
DOI:10.1186/s42400-025-00481-3delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Malware family classification is crucial for threat detection, yet existing methods struggle with generalization, multi-task adaptability, and interpretability. We propose VIMAR, a unified vision–language model that supports classification, similarity detection, and open-world analysis via explanation-rich supervision and a two-stage training pipeline. On the Malimg dataset, VIMAR achieves 94.2% accuracy in family classification, surpassing the best CNN baseline by +3.1%. It also attains 85.2% and 88.0% accuracy in zero-shot and few-shot settings, significantly outperforming vision–language baselines. Moreover, its reasoning outputs align well with human judgments. The codebase and scripts will be released to the community.
Keywords:
Malware family classification
Vision-language model
Zero-shot learning
Deep learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

C
Cybersecurity
IF:
3.7
Papers:
579
Citations:
1.0K

Organization

C
College of Cyberspace Security
Scholars:
14
Papers: 7
Citations: 0