arXiv Computer Vision

Difficulty-Aware Sample Allocation for Adaptive Data Augmentation in Semantic Segmentation

The paper proposes Difficulty-Aware Sample Allocation (DASA), a framework that assigns stronger data augmentation to training samples deemed more difficult based on a composite difficulty score. This score integrates prediction ambiguity, training loss, class rarity, and boundary complexity, and is used to modulate augmentation strength during iterative training. Experiments on Oxford‑IIIT Pet and binary Pascal VOC with U‑Net, DeepLabV3, and SegFormer‑B0 demonstrate that DASA outperforms standard training and matches or exceeds single‑signal adaptive baselines, notably raising DeepLabV3’s mIoU from 0.633 to 0.740 on Oxford‑IIIT Pet.

arXiv Machine Learning
Aug 4

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

arXiv:2604. 15622v3 Announce Type: replace-cross Abstract: Always-on contextual AI runs language-aligned vision foundation models (VFMs) on edge devices, where the on-device model is the dominant continuous compute cost under strict latency and power limits.

By Yiwei Zhao, Yi Zheng, Huapeng Su, Jieyu Lin, Stefano Ambrogio, Cijo Jose, Michael Ramamonjisoa, Patrick Labatut, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li
arXiv Computer Vision
Sep 17

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.

By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa