Return
Enhancing AI Binocular Vision for Camellia oleifera Fruit Localization Using Zero-Shot Lightweight Segmentation Model and 3D Surface Optimization
S
L
L
M
H
DOI:10.1016/j.eng.2026.06.020.png)
Abstract
En 中文
The segmentation and three-dimensional (3D) localization accuracy of Camellia oleifera fruits by harvesting robots are critical factors influencing their success rate. To address the challenges of high annotation costs and significant localization errors in fruit segmentation and localization, an artificial intelligence (AI)-based binocular vision method for Camellia oleifera fruit localization is proposed using zero-shot lightweight segmentation and 3D surface optimization models. First, the large vision model Grounding DINO generates the bounding boxes of fruit objects using the text prompt “fruit”. These bounding boxes were provided as prompts to the segment anything model (SAM) to obtain segmentation masks, eliminating the need for manual annotation. To meet the real-time application requirements of embedded devices, a knowledge transfer method is proposed. The segmentation results of the large models were used as labels to train YOLO11n-seg, a considerably small model without requiring manual annotation, achieving a zero-shot annotation approach. Moreover, model pruning and knowledge distillation were applied to YOLO11n-seg, compressing model parameters and improving runtime efficiency. The results demonstrated that the pruned and distilled YOLO11n-seg model achieved a mean intersection over union (mIoU) of 83.52% and F1 score of 90.87%, improving 1.46% and 0.97% over the original model, respectively. Although its performance is less than that of the Grounding DINO + SAM, the simplified YOLO11n-seg model features a parameter count of only 2.06 million (M), well suited for deployment on embedded devices. Compared with YOLOv8s-seg, YOLOv9c-seg, and YOLO11m-seg, it requires only 16.98%, 4.57%, and 5.85% of their computational costs, respectively, while achieving an average inference speed of 108.20 frames per second (FPS). Finally, the proposed mask-based 3D localization of fruits achieved a localization error of 9.94 mm, with 3.79 mm reduction compared with the bounding box center-based method. Overall, the proposed method provides an efficient solution for the recognition and localization of Camellia oleifera fruits and other similar types of fruits.
Keywords:
Zero-shot learning
Large vision model
Instance segmentation
Camellia oleifera fruit
Fruit location
Journal
IF:
11.6
Papers:
2.7K
Citations:
1.5W
