Return
A visual transformer and multi-task learning framework for building generalisation from raster maps
DOI:10.1080/17538947.2026.2724180.png)
Abstract
En 中文
Building generalisation transforms detailed large-scale building representations into suitable simplified forms for smaller-scale urban topographic maps. Existing deep learning-based approaches for raster building maps are largely built on convolutional neural networks (CNNs) or generative adversarial networks (GANs). However, their local inductive biases or limited global spatial awareness often cause geometric distortions in buildings. Although large-vision foundation models excel in holistic modeling, their application has focused mainly on building extraction. To address these issues, SAM-BG, a Building Generalisation framework that integrates the Segment Anything Model with multitask learning, was proposed. Specifically, low-rank adaptation was employed to efficiently fine-tune the vision transformer (ViT) encoder and transfer global representation capability to the cartographic domain. Subsequently, a geometry-aware multitask architecture with explicit boundary supervision was introduced to alleviate contour distortion. Moreover, a Boundary Feature Enhancement Module (BFEM) that uses dual-source-guided gating and multi-scale feature pyramids was designed to reinforce boundary responses. Multi-scale experiments on synthetic and real-world datasets showed that SAM-BG outperformed benchmark methods on boundary-sensitive metrics of buildings. This study presents a viable end-to-end solution for building generalisation and validates the potential of vision foundation models within domain-specific context of cartography.
Keywords:
Building generalisation
raster maps
vision transformer
low-rank adaptation
multi-task learning
Journal
IF:
4.9
Papers:
1.9K
Citations:
4.7K

