Return
Recognizing temporary construction site objects using CLIP-based few-shot learning and multi-modal prototypes
DOI:10.1016/j.autcon.2024.105542.png)
Abstract
En 中文
Visual understanding of temporary on-site objects is essential for robots and project management in construction. Implementation of deep learning algorithms is challenging on construction sites due to high data annotation cost, demanding computational power, and lack of large-scale training datasets. Recognizing on-site temporary objects demands the algorithms to learn in a data-efficient way. To fill this gap, a Contrastive Language-Image Pretraining (CLIP)-based few-shot learning algorithm to recognize temporary objects with limited image samples is proposed. The study builds ImageNet-based Similarity Cache with inter-class similarity distribution. The proposed algorithm is evaluated on a newly created TOCS dataset and on the public SODA dataset. Compared with CLIP zero-shot algorithm, the classification accuracy improves from 23.17% to 73.09% with 16-shot learning on SODA, and from 58.33% to 83.33% with 1-shot learning on TOCS. The study indicates that fewshot learning with vision language models (VLM) is promising to improve visual intelligence on construction sites.
Keywords:
Construction sites
CLIP
Few -shot image classification (FSIC)
Computer vision
Robots
Journal
IF:
11.5
Papers:
6.2K
Citations:
4.2W

