Return
DepthPlay-Agent: Agentic Learning and Fusion for Depth Estimation
DOI:10.1109/OJSP.2026.3703724.png)
Abstract
En 中文
Monocular depth estimation in complex, dynamic environments remains challenging due to rapid object motion, texture repetitions, occlusions, and strong geometric constraints inherent in structured scenes. To address these challenges, we propose DepthPlay-Agent (DEAN), an agentic multi-model fusion framework that builds upon SOTA depth estimator backbones such as UNet++, HybridDepth, DINOv3, and ZoeDepth under a learnable Mixture-of-Experts controller. This controller embeds scene context through transformers and uses task priors to dynamically assign pe-pixel fusion weights. A domain-aware refinement module further enforces geometric consistency using planar and semantic segmentation cues. Beyond static fusion, DEAN introduces an agentic inference layer that dynamically regulates expert contributions and refinement strategies, enabling adaptive and interpretable decision-making. Experiments across four benchmarks (SoccerNet-Depth, KITTI, MPI-Sintel, NYU Depth V2) demonstrate that DEAN consistently achieves SOTA performance, improving AbsRel by up to 5% and reducing SILog by 8% over strong baselines. By coupling multi-modal intelligence with structured geometric reasoning, DEAN establishes a new paradigm for adaptive, context-aware depth estimation in dynamic real-world domains.
Keywords:
Monocular depth estimation
agentic learning
multimodal learning
decision making
Journal
IF:
2.7
Papers:
140
Citations:
535

