arrow
Return

ROARS: A Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis

delete2026-08-05
delete0
PRE
AI
J
Jingrui Zhang
梁风 cover
梁风 (Feng Liang)
Y
Yong Zhang
Z
Zixuan Shangguan
Z
Zhida Li
X
Xiaoyi Fan
V
Victor C. M. Leung
G
Guanbin Li
X
Xiping Hu
DOI:10.1109/tcc.2026.3721293delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The paradigm shift of remote sensing (RS) analysis toward cloud-based large-scale computing demands intelligent methods capable of interpreting petabyte-scale, high-resolution heterogeneous imagery. Multimodal Large Language Models (MLLMs) have emerged as a transformative solution for such data-intensive tasks, offering unprecedented capabilities in open-set perception and high-level semantic reasoning. However, directly applying general MLLM architectures to the RS domain faces significant challenges: the constraint of fixed-resolution vision encoders causes severe information loss for small objects, while the reliance on single-layer feature fusion neglects critical spatial cues preserved in intermediate layers. To address these challenges, we propose ROARS: a Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis, tailored for high-fidelity interpretation. First, we introduce a Native-Resolution Vision Encoder tailored to process images of arbitrary aspect ratios and scales without downsampling, thereby strictly preserving pixel-level details. Second, we design a Shortcut-Based Fusion Mechanism that bridges the semantic gap by injecting multi-level visual features into the language model. Furthermore, to enhance global contextual understanding, we incorporate a Multi-Granularity Token Pooling strategy that provides fine- and coarse-grained visual representations. Extensive experiments on the VRSBench demonstrate that our framework significantly outperforms state-of-the-art approaches, particularly in fine-grained object perception and complex spatial reasoning. These results validate the proposed method as a robust and efficient core component for next-generation cloud-based RS interpretation systems.
Keywords:
Remote Sensing
Native Resolution
Multimodal Large Language Model
Vision-Language Information Fusion

Journal

I
IEEE Transactions on Cloud Computing
IF:
5
Papers:
1.8K
Citations:
4.3K

Organization

S
sun yat-sen university
Scholars:
1.3K
Papers: 382
Citations: 0
S
Shenzhen MSU-BIT University
Scholars:
72
Papers: 34
Citations: 0
Cited Papers

Cited Papers

No cited papers available