Return
Light Field Referring Segmentation: A Benchmark and an LLM-Based Approach
S
Q
G
T
Y
Q
DOI:10.1109/TBC.2026.3659013.png)
Abstract
En 中文
Referring image segmentation (RIS) is a challenging task that requires models to segment objects based on natural language descriptions. Existing RIS models have limited leverage of geometric information, resulting in multimodal mismatch between language and vision. In contrast to conventional 2D images, light field imaging gathers rays emitted from light sources in all directions. This unique characteristic enriches the comprehensive understanding of scenes, which provides us with a new way to optimize RIS. In this paper, we propose the first light field referring segmentation dataset, which contains rich occluded objects and depth-referring descriptions. Afterward, we benchmark the performance of existing 2D referring image segmentation methods on the proposed dataset. The results revealed that these methods show limited efficacy in occluded scenes and depth-based descriptions of scenes. To address this issue, we propose a novel framework, termed LFLLM, for light field referring segmentation. Specifically, we propose a Center Angular Aggregation Module that warps the views adjacent to the central view to prevent feature occlusion caused by viewpoints misalignment, and a Depth Convergence Module that adds a depth token into the LLMs to leverage the depth information in the light field. Extensive experiments demonstrate that our approach outperforms the current state-of-the-art methods. The dataset and code are available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/ShishunTian/LFLLM-TBC2026</uri>
Keywords:
Referring image segmentation
light field
dataset
large language models
Journal
IF:
4.8
Papers:
2.1K
Citations:
3.0K
