The paper introduces Recursive Block-Diagonal Coupling (RBDC), a training protocol that builds wide vision models by recursively coupling narrower, independently trained models in a parameter‑free block‑diagonal manner. RBDC allows flexible allocation of training budgets across all models and, when applied to vision transformers (DeiT) and convolutional networks (ResNet) on ImageNet, achieves a 30% reduction in FLOPs while maintaining similar test accuracies. Additionally, models trained with RBDC outperform those from existing growth methods at the same training FLOPs and serve as stronger backbones for downstream tasks such as object detection and instance segmentation.
By Maxim Henry, Adrien Deli\`ege, S\'ebastien Pi\'erard, Marc Van Droogenbroeck
arXiv:2609.35490v2 Announce Type: replace
Abstract: Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning....
By Mohammad Mahdi, Nedyalko Prisadnikov, Yuqian Fu, Carmelo Scribano, Danda Pani Paudel, Luc Van Gool
arXiv:2609.36875v1 Announce Type: new
Abstract: Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Any...
By Jingdong Zhang, Xin Li, Jan Kautz, Wenping Wang, Chris Choy
The paper introduces OVRSISBench, a unified benchmark for open‑vocabulary remote sensing image segmentation, and evaluates existing OVS/OVRSIS models, uncovering their shortcomings in remote sensing contexts. Leveraging insights from this evaluation, the authors propose RSKT‑Seg, a new framework featuring a Multi‑Directional Cost Map Aggregation module, an Efficient Cost Map Fusion transformer, and a Remote Sensing Knowledge Transfer module. Experiments on the benchmark demonstrate that RSKT‑Seg outperforms strong baselines by +3.8 mIoU and +5.9 mACC while achieving twice the inference speed.
By Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li
DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.
By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.
By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa