Return
VIMAR: vision-language informed malware analysis and reasoning model
DOI:10.1186/s42400-025-00481-3.png)
Abstract
En 中文
Malware family classification is crucial for threat detection, yet existing methods struggle with generalization, multi-task adaptability, and interpretability. We propose VIMAR, a unified vision–language model that supports classification, similarity detection, and open-world analysis via explanation-rich supervision and a two-stage training pipeline. On the Malimg dataset, VIMAR achieves 94.2% accuracy in family classification, surpassing the best CNN baseline by +3.1%. It also attains 85.2% and 88.0% accuracy in zero-shot and few-shot settings, significantly outperforming vision–language baselines. Moreover, its reasoning outputs align well with human judgments. The codebase and scripts will be released to the community.
Keywords:
Malware family classification
Vision-language model
Zero-shot learning
Deep learning
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

