Return
Neuro-Symbolic Vision Framework for Construction: Knowledge Graph-Augmented Monitoring and Productivity Analysis of Precast Concrete Assembly
J
E
J
T
DOI:10.1093/jcde/qwag067.png)
Abstract
En 中文
Monitoring precast concrete (PC) assembly is hindered by logical inconsistencies, occlusions, and the inability of frame-by-frame analysis to capture temporal sequences. To address this, we propose a neuro-symbolic vision framework integrating data-driven deep learning with knowledge-based reasoning. The system combines YOLOv11, DeepSORT, and Video Masked Autoencoders (VideoMAE), enhanced by a knowledge-graph reasoning engine that applies spatial and temporal domain constraints. Methodologically, vision outputs are temporally aligned and then passed through a multi-stage graph inference process that performs sequence and prerequisite filtering, fuses four complementary confidence scores, and continuously adapts its temporal parameters from site-specific observations; the framework was trained and evaluated on a multimodal dataset of 12,000 annotated images and over 12,000 activity clips. Experimental validation on real-world videos demonstrates that this integrated approach achieves 96.8% precision, 96.2% recall, and 96.5% F1 score while sustaining real-time throughput (28.3 FPS), significantly outperforming standalone VideoMAE by 8.3 percentage points. Furthermore, productivity analysis of 142 components shows minimal error (Mean Absolute Percentage Error (MAPE): 1.3–3.4%) and high correlation (τ ≥ 0.90) with actual data. These results confirm that neuro-symbolic integration effectively overcomes pure computer vision limitations, enabling accurate, real-time monitoring to improve project efficiency and decision-making in PC construction.
Journal
IF:
6.1
Papers:
392
Citations:
3.2K
