arXiv Computer Vision

GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation

arXiv Computer Vision
Sep 2

C-DiffSET: Leveraging Latent Diffusion for SAR-to-EO Image Translation with Confidence-Guided Reliable Object Generation

C‑DiffSET is a SAR‑to‑EO image translation framework that uses a pretrained Latent Diffusion Model to adapt SAR imagery to the EO domain. The method exploits the pretrained VAE encoder’s ability to map SAR and EO images into a shared latent space, even when SAR inputs contain varying noise levels. A confidence‑guided diffusion loss further improves pixel‑wise fidelity by reducing artifacts such as appearing or disappearing objects, leading to state‑of‑the‑art results across multiple datasets.

By Jeonghyeok Do, Jaehyup Lee, Munchurl Kim
arXiv Machine Learning
Sep 10

Cross-modal learning for SAR target recognition using optical vision foundation models

The paper proposes a cross‑modal framework that uses a frozen DINOv3 optical vision foundation model to create class‑level prototypes for Synthetic Aperture Radar (SAR) target recognition. By aligning SAR embeddings to these optical prototypes, the SAR model learns to classify SAR images without needing paired optical data. Experiments on the UNICORNv2 dataset show that this prototype alignment improves SAR classification accuracy compared to baseline methods and yields clearer class separation in the embedding space.

By Lucas Hirsch, James R. Hopgood, Javid Khan, Yoann Altmann, Mike E. Davies
arXiv Computer Vision
Sep 7

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

The paper introduces BMT, a unified hierarchical Vision Transformer that jointly performs SAR-to-optical image translation and semantic segmentation. It incorporates a LocalViTBlock, an enhanced output module, a ControlNet-style conditional injection, and a bounded Kendall uncertainty weighting scheme to balance the two tasks. Experiments on paired and unpaired datasets demonstrate competitive performance in both translation quality and segmentation accuracy.

By Siyuan Liu, Xuze Zhang, Yongshun Wang, Licong Pan, Hang Liu, Huihui Li
Hugging Face Trending Papers
Sep 2

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

The paper demonstrates a lightweight method to adapt general-purpose vision‑language models (VLMs) for multispectral and synthetic aperture radar (SAR) image understanding. By rendering each observation as five optical views and one SAR view, naming them in the prompt, and applying LoRA to the language network and selected visual transformer blocks, the authors enable VLMs to process band composites, spectral indices, and radar backscatter without retraining a new foundation model. On a balanced six‑class land‑cover benchmark from BigEarthNet‑v2, the adapted Qwen3‑VL achieves a micro F1 score of 0.8275, and the same protocol improves four other VLMs and transfers to flood verification and captioning tasks.

arXiv Computer Vision
Sep 3

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

The paper presents a lightweight method to adapt general‑purpose vision‑language models (VLMs) for multispectral and synthetic aperture radar (SAR) image understanding. By rendering each observation as five optical views and one SAR view, naming them in the prompt, and applying LoRA to the language network and selected visual transformer blocks, the authors enable VLMs to process band composites, spectral indices, and radar backscatter without retraining a new foundation model. On a balanced six‑class land‑cover benchmark from BigEarthNet‑v2, the adapted Qwen3‑VL achieves a micro F1 of 0.8275, and the same protocol improves four other VLMs and transfers to flood verification and captioning tasks. "whyItMatters":"The study shows that existing VLMs can be repurposed for multispectral and SAR tasks through simple input rendering and compact LoRA adaptation, avoiding the need for dedicated encoders and domain pretraining."

By Shanji Liu, Kelu Yao, Junxiao Xue, Chenghui Lv, Xiangyang Miao, Yekai Huang, Yaying Chen, Chao Li
arXiv AI
Jun 6

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

arXiv:2602. 19190v4 Announce Type: replace-cross Abstract: Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applications.

By Xiaokun Zhang, Yi Yang, Ziqi Ye, Baiyun, Xiaorong Guo, Qingchen Fang, Ruyi Zhang, Xinpeng Zhou, Haipeng Wang
arXiv Computer Vision
Sep 25

OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding

OptiSAR-Net++ introduces a new cross‑domain remote sensing visual grounding task (CD‑RSVG) and the first large‑scale benchmark dataset, OptSAR‑RSVG. The framework replaces Transformer decoding with a CLIP‑based contrastive approach, employing a patch‑level Low‑Rank Adaptation Mixture of Experts for efficient cross‑domain feature decoupling and a text‑guided dual‑gate fusion module for improved semantic‑visual alignment. Experiments show state‑of‑the‑art performance on OptSAR‑RSVG and DIOR‑RSVG, with notable gains in localization accuracy and computational efficiency.

By Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan