arrow
Return

Binary Code Representation With Well-Balanced Instruction Normalization

delete2023-01-01
delete3
delete
OA
AI
H
Hyungjoon Koo *
S
Soyeon Park
D
Daejin Choi
T
Taesoo Kim
DOI:10.1109/ACCESS.2023.3259481delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The recovery of contextual meanings on a machine code is required by a wide range of binary analysis applications, such as bug discovery, malware analysis, and code clone detection. To accomplish this, advancements on binary code analysis borrow the techniques from natural language processing to automatically infer the underlying semantics of a binary, rather than replying on manual analysis. One of crucial pipelines in this process is instruction normalization, which helps to reduce the number of tokens and to avoid an out-of-vocabulary (OOV) problem. However, existing approaches often substitutes the operand(s) of an instruction with a common token (e.g., callee target ? FOO), inevitably resulting in the loss of important information. In this paper, we introduce well-balanced instruction normalization (WIN), a novel approach that retains rich code information while minimizing the downsides of code normalization. With large swaths of binary code, our finding shows that the instruction distribution follows Zipf's Law like a natural language, a function conveys contextually meaningful information, and the same instruction at different positions may require diverse code representations. To show the effectiveness of WIN, we present DeepSemantic that harnesses the BERT architecture with two training phases: pre-training for generic assembly code representation, and fine-tuning for building a model tailored to a specialized task. We define a downstream task of binary code similarity detection, which requires underlying code semantics. Our experimental results show that our binary similarity model with WIN outperforms two state-of-the-art binary similarity tools, DeepBinDiff and SAFE, with an average improvement of 49.8% and 15.8%, respectively.
Keywords:
Task analysis
Binary codes
Bit error rate
Semantics
Computer architecture
Computational modeling
Transformers
Binary code
code representation
BERT
well-balanced instruction normalization
binary code similarity detection

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

G
Georgia Institute of Technology
Scholars:
1.8W
Papers: 1.4W
Citations: 5.9W
S
sungkyunkwan university (skku)
Scholars:
3.7W
Papers: 3.6W
Citations: 49
U
university system of georgia
Scholars:
7.3W
Papers: 6.5W
Citations: 101
researcher View more organizations