Return
Fish mass estimation based on monocular depth estimation and instance segmentation
DOI:10.1016/j.engappai.2025.113513.png)
Abstract
En 中文
Accurate and non-invasive fish body mass estimation is essential for intelligent aquaculture. Existing monocular vision methods struggle to estimate the mass of freely swimming fish, while stereo vision is sensitive to environmental interference and computationally complex. To address these challenges, this study proposes a mass estimation method leveraging cross-modal cues from monocular depth estimation and instance segmentation, optimized for turbid water and occlusion in aquaculture scenarios. For depth estimation, a Dense Perception Enhanced Mixed Vision Transformer network is designed to fuse multi-scale features and enhance local details in blurry regions, while a gated Top-K unit suppresses low-confidence depth features and occlusion interference, reducing the absolute relative error by approximately 12.6%. For instance segmentation, a dual-branch network improves complementary perception of blurred features, and a spatial-enhanced dynamic head module achieves high-precision segmentation of densely packed fish, improving the mean average precision by approximately 2.76%. Experimental results demonstrate that the proposed method achieves superior performance compared to existing methods in both depth estimation and instance segmentation. In real aquaculture scenarios, it attains a mean absolute error of 0.0172, root mean square error of 0.0211, and coefficient of determination of 0.9624, validating its effectiveness as a reliable automated solution for intelligent aquaculture.
Journal
IF:
8
Papers:
5.3K
Citations:
3.5W

