Hugging Face Trending Papers

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models

The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision models, particularly recent ViT-based foundation models in dense prediction tasks. Instance segmentation, a typical dense visual prediction task in the remote sensing field, faces similar challenges.

arXiv Computer Vision
Sep 21

Recursive Block-Diagonal Coupling for Resource-Efficient Training of Vision Models

The paper introduces Recursive Block-Diagonal Coupling (RBDC), a training protocol that builds wide vision models by recursively coupling narrower, independently trained models in a parameter‑free block‑diagonal manner. RBDC allows flexible allocation of training budgets across all models and, when applied to vision transformers (DeiT) and convolutional networks (ResNet) on ImageNet, achieves a 30% reduction in FLOPs while maintaining similar test accuracies. Additionally, models trained with RBDC outperform those from existing growth methods at the same training FLOPs and serve as stronger backbones for downstream tasks such as object detection and instance segmentation.

By Maxim Henry, Adrien Deli\`ege, S\'ebastien Pi\'erard, Marc Van Droogenbroeck
arXiv AI
Aug 19

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

The paper introduces OVRSISBench, a unified benchmark for open‑vocabulary remote sensing image segmentation, and evaluates existing OVS/OVRSIS models, uncovering their shortcomings in remote sensing contexts. Leveraging insights from this evaluation, the authors propose RSKT‑Seg, a new framework featuring a Multi‑Directional Cost Map Aggregation module, an Efficient Cost Map Fusion transformer, and a Remote Sensing Knowledge Transfer module. Experiments on the benchmark demonstrate that RSKT‑Seg outperforms strong baselines by +3.8 mIoU and +5.9 mACC while achieving twice the inference speed.

By Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li
arXiv Computer Vision
3d ago

DiDA: Video Object Segmentation with Distillation Learning of Deformable Attention

DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.

By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
arXiv Computer Vision
Sep 17

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.

By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
arXiv Computer Vision
Sep 24

A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing

The paper introduces O$^2$-VG, a unified framework for oriented object visual grounding in remote sensing images, comprising three complementary models: O$^2$-VG-Trans, a cross‑modality transformer; O$^2$-VG-Uni, which predicts universal oriented proposals; and O$^2$-VG-VLM, an autoregressive vision‑language model that generates oriented bounding boxes. It also presents DIOR‑R‑SVG, a new dataset containing image, expression, and oriented box triplets for training and evaluation. The framework demonstrates superior performance across multiple benchmarks and is supported by publicly available code.

By Zeyu Ding, Yong Zhou, Jiaqi Zhao, Wen-Liang Du, Xixi Li, Hancheng Zhu, Rui Yao, Abdulmotaleb El Saddik
arXiv Machine Learning
Jul 20

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

arXiv:2607. 15942v1 Announce Type: cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks.

By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")