
3D GeoInfo & SDSC 2025
20th 3D GeoInfo Conference | 9th Smart Data and Smart Cities Conference
02 - 05 September 2025 | Kashiwa Campus, University of Tokyo, Japan
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
Session 12-b: SDSC - Transportation Location: Media Hall / Kashiwa Library Session Chair: Takahiro Yoshida | |
| Presentation 4 | |
VLM-Based Building Change Detection with CNN-Transformer 1: University of Osnabrueck, Germany; 2: German Aerospace Center (DLR), Germany Rapid urbanization and environmental changes necessitate precise and scalable methods for detecting building changes in satellite imagery, critical for urban planning, disaster response, and smart city analytics. Traditional approaches, relying on handcrafted features or shallow learning algorithms, often fail to capture the complex spatio-temporal patterns inherent in high-resolution remote sensing data. Recent advancements in deep learning, particularly convolutional neural networks (CNNs) and Transformer architectures, have significantly improved the ability to model these intricate relationships, enabling more accurate detection of structural changes (Chen et al., 2022). However, these methods typically require extensive labeled datasets and domain-specific fine-tuning, limiting their scalability and adaptability across diverse scenarios. The emergence of vision-language models (VLMs), which integrate textual and visual information, offers a promising avenue to enhance contextual understanding and focus on task-relevant features (Tao et al., 2025). Despite these advances, adapting pretrained VLMs to remote sensing tasks, such as building change detection, remains a significant challenge due to differences between natural imagery and satellite data. In this study we propose a novel framework that combines Grounding DINO (Liu et al., 2023), a pretrained VLM, with a hybrid CNN-Transformer architecture. Our approach leverages text-informed preprocessing to generate building masks, which guide a lightweight ResNet18 (He et al., 2016) backbone and a custom Transformer encoder to focus on structural and spectral changes. By amplifying building-related features and employing a combined loss function, our framework achieves promising performance on a benchmark change detection dataset. The proposed method minimizes reliance on extensive finetuning, offering a scalable solution for urban monitoring and Earth observation applications. | |