arXiv AI

SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

arXiv:2606. 20523v1 Announce Type: cross Abstract: Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture radar (SAR) remain limited.

arXiv Machine Learning
Sep 10

Cross-modal learning for SAR target recognition using optical vision foundation models

The paper proposes a cross‑modal framework that uses a frozen DINOv3 optical vision foundation model to create class‑level prototypes for Synthetic Aperture Radar (SAR) target recognition. By aligning SAR embeddings to these optical prototypes, the SAR model learns to classify SAR images without needing paired optical data. Experiments on the UNICORNv2 dataset show that this prototype alignment improves SAR classification accuracy compared to baseline methods and yields clearer class separation in the embedding space.

By Lucas Hirsch, James R. Hopgood, Javid Khan, Yoann Altmann, Mike E. Davies
arXiv AI
Jun 6

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

arXiv:2602. 19190v4 Announce Type: replace-cross Abstract: Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applications.

By Xiaokun Zhang, Yi Yang, Ziqi Ye, Baiyun, Xiaorong Guo, Qingchen Fang, Ruyi Zhang, Xinpeng Zhou, Haipeng Wang
Hugging Face Trending Papers
Sep 2

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

The paper demonstrates a lightweight method to adapt general-purpose vision‑language models (VLMs) for multispectral and synthetic aperture radar (SAR) image understanding. By rendering each observation as five optical views and one SAR view, naming them in the prompt, and applying LoRA to the language network and selected visual transformer blocks, the authors enable VLMs to process band composites, spectral indices, and radar backscatter without retraining a new foundation model. On a balanced six‑class land‑cover benchmark from BigEarthNet‑v2, the adapted Qwen3‑VL achieves a micro F1 score of 0.8275, and the same protocol improves four other VLMs and transfers to flood verification and captioning tasks.

arXiv Computer Vision
Sep 3

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

The paper presents a lightweight method to adapt general‑purpose vision‑language models (VLMs) for multispectral and synthetic aperture radar (SAR) image understanding. By rendering each observation as five optical views and one SAR view, naming them in the prompt, and applying LoRA to the language network and selected visual transformer blocks, the authors enable VLMs to process band composites, spectral indices, and radar backscatter without retraining a new foundation model. On a balanced six‑class land‑cover benchmark from BigEarthNet‑v2, the adapted Qwen3‑VL achieves a micro F1 of 0.8275, and the same protocol improves four other VLMs and transfers to flood verification and captioning tasks. "whyItMatters":"The study shows that existing VLMs can be repurposed for multispectral and SAR tasks through simple input rendering and compact LoRA adaptation, avoiding the need for dedicated encoders and domain pretraining."

By Shanji Liu, Kelu Yao, Junxiao Xue, Chenghui Lv, Xiangyang Miao, Yekai Huang, Yaying Chen, Chao Li
arXiv Computer Vision
Sep 25

OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding

OptiSAR-Net++ introduces a new cross‑domain remote sensing visual grounding task (CD‑RSVG) and the first large‑scale benchmark dataset, OptSAR‑RSVG. The framework replaces Transformer decoding with a CLIP‑based contrastive approach, employing a patch‑level Low‑Rank Adaptation Mixture of Experts for efficient cross‑domain feature decoupling and a text‑guided dual‑gate fusion module for improved semantic‑visual alignment. Experiments show state‑of‑the‑art performance on OptSAR‑RSVG and DIOR‑RSVG, with notable gains in localization accuracy and computational efficiency.

By Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan
arXiv AI
Sep 15

Multimodal-Multiresolution Foundation Model for Lunar Remote Sensing

The paper introduces a multimodal foundation model for lunar remote sensing, trained from scratch on SomBench—a dataset of nearly two million co‑registered tile bundles across 11 modalities at 1 m and 100 m resolutions. The model extends the TerraMind masked‑token architecture with lunar‑specific features such as explicit acquisition geometry and joint training of two spatial scales, and employs FlexiViT patch embeddings for adaptable patch sizes. Evaluation on crater detection, irregular mare patch segmentation, and polar ice prospectivity regression shows that the pretrained model matches or surpasses ImageNet‑pretrained baselines, with notable label efficiency and effective adaptation via LoRA.

By Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal, Rachel A. Slank, Geoffrey Dawson, Hiyam Debary, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Nikolaos Dionelis, Ankur Kumar, Campbell D. Watson, Manil Maskey, Rebekah I. Dawson-Rigas, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
arXiv Computer Vision
Sep 7

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

The paper introduces BMT, a unified hierarchical Vision Transformer that jointly performs SAR-to-optical image translation and semantic segmentation. It incorporates a LocalViTBlock, an enhanced output module, a ControlNet-style conditional injection, and a bounded Kendall uncertainty weighting scheme to balance the two tasks. Experiments on paired and unpaired datasets demonstrate competitive performance in both translation quality and segmentation accuracy.

By Siyuan Liu, Xuze Zhang, Yongshun Wang, Licong Pan, Hang Liu, Huihui Li