arrow
Return

Recognizing temporary construction site objects using CLIP-based few-shot learning and multi-modal prototypes

delete2024-09-01
delete0
PRE
AI
Y
Yuanchang Liang *
P
Prahlad Vadakkepat
D
David Kim Huat Chua
S
Shuyi Wang
李志刚 (Zhigang Li)
张书香 (Shuxiang Zhang)
DOI:10.1016/j.autcon.2024.105542delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual understanding of temporary on-site objects is essential for robots and project management in construction. Implementation of deep learning algorithms is challenging on construction sites due to high data annotation cost, demanding computational power, and lack of large-scale training datasets. Recognizing on-site temporary objects demands the algorithms to learn in a data-efficient way. To fill this gap, a Contrastive Language-Image Pretraining (CLIP)-based few-shot learning algorithm to recognize temporary objects with limited image samples is proposed. The study builds ImageNet-based Similarity Cache with inter-class similarity distribution. The proposed algorithm is evaluated on a newly created TOCS dataset and on the public SODA dataset. Compared with CLIP zero-shot algorithm, the classification accuracy improves from 23.17% to 73.09% with 16-shot learning on SODA, and from 58.33% to 83.33% with 1-shot learning on TOCS. The study indicates that fewshot learning with vision language models (VLM) is promising to improve visual intelligence on construction sites.
Keywords:
Construction sites
CLIP
Few -shot image classification (FSIC)
Computer vision
Robots

Journal

Automation in Construction cover
Automation in Construction
IF:
11.5
Papers:
6.2K
Citations:
4.2W

Organization

N
National University of Singapore
Scholars:
7.5W
Papers: 6.5W
Citations: 11.4W