Enhancing Inter-procedural Static Analysis with Selective LLM-Driven Data-flow Summarization
Abstract
In this work, we propose LLSUM, a novel approach that selectively integrates LLMs to summarize complex methods that static analysis cannot handle effectively, seamlessly incorporating the results back into the analysis. We introduce a conditional summary representation that bridges LLM-generated summaries with precise symbolic representations, an analysis algorithm to gather sufficient context for generating reusable summaries, and a scheduling strategy to identify when and where LLMs are most beneficial for improving efficiency and precision.
We evaluate LLSUM on several popular open-source Java projects, achieving a 6%-43% accuracy improvement and up to a 16× speed enhancement over baseline methods. Additionally, LLSUM uncovered 28 unique zero-day vulnerabilities in real-world applications, 7 of which have received CVE identifiers. These results demonstrate the effectiveness of LLSUM in improving the precision and efficiency of static taint analysis, paving the way for more robust and scalable vulnerability detection.

