Return
GUIGroup: Enabling Functional Layout Grouping With Multimodal Large Language Models
Y
Y
Z
X
刘
Z
X
鲍
DOI:10.1109/tse.2026.3688725.png)
Abstract
En 中文
Modern graphical user interfaces usually contain a large number of scattered GUI widgets, which need to be organized into functionally related layout groups to provide a good user experience and design consistency. As the core task of GUI automated analysis, functional layout grouping can not only support interface design evaluation and optimization, but also provide an important foundation for downstream tasks such as GUI to code generation, interface understanding and automated testing. Previous functional layout grouping research mainly relies on traditional computer vision methods to group interface elements through heuristic rules and psychological principles. Unfortunately, due to the lack of adaptability of rule-based methods and insufficient understanding ability in dense layouts, the performance of these methods is not satisfactory. To address these limitations, we propose the <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">GUIGroup</small> method to achieve accurate functional layout grouping by identifying GUI widgets in screenshots and combining widget detection techniques with MLLM. The evaluation on a dataset of 1,612 GUIs manually collected from 596 Android apps shows that <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">GUIGroup</small> significantly outperforms the baseline method with an F1 score of 0.5353, which is 12.65% higher than the best baseline. In addition, ablation studies deeply verify the key contributions of each step, and user studies further confirm the superiority and practical value of the method in real application scenarios.
Keywords:
GUI
MLLM
functional layout grouping
Journal
IF:
5.6
Papers:
2.8K
Citations:
1.1W
