arXiv AI

Parameter-Efficient Fine-Tuning of Large Pretrained Models for Instance Segmentation Tasks

arXiv:2606. 01947v1 Announce Type: cross Abstract: Research and applications in artificial intelligence have recently shifted with the rise of large pretrained models, which deliver state-of-the-art results across numerous tasks.

Hugging Face Trending Papers
Jun 24

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models

The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision models, particularly recent ViT-based foundation models in dense prediction tasks. Instance segmentation, a typical dense visual prediction task in the remote sensing field, faces similar challenges.

arXiv Machine Learning
Jun 18

Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models

arXiv:2509. 22020v2 Announce Type: replace Abstract: While recent advances in machine learning have equipped Weather Foundation Models (WFMs) with substantial generalization capabilities across diverse downstream tasks, the escalating computational requirements associated with their expanding scale increasingly hinder practical deployment.

By Shilei Cao, Hehai Lin, Jiashun Cheng, Yang Liu, Guowen Li, Xuehe Wang, Juepeng Zheng, Haoyuan Liang, Meng Jin, Chengwei Qin, Hong Cheng, Haohuan Fu
arXiv Computer Vision
Aug 24

ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation

ES‑VP introduces Energy‑Shaped Visual Prompting, a method that generates image‑specific prompts through low‑rank initialization and energy‑guided dynamic adaptation. It achieves higher performance than existing single‑prompt and diverse‑prompt approaches while using far fewer parameters. Experiments on five architectures and fifteen datasets show consistent superiority, including a 2.6% accuracy gain over DAM‑VP on CLIP with 590× fewer prompt parameters.

By Can Jin, Ying Li, Jingchen Sun, Hongwu Peng, Jiahui Zhao, Yang Zhou, Lei Li, Dimitris N. Metaxas
arXiv Computer Vision
Sep 17

Position Anchor Tuning: Towards Efficient Adaptation of Pre-Trained Point Cloud Transformers

Position Anchor Tuning (PAT) is a parameter‑efficient fine‑tuning method for pre‑trained point cloud transformers that improves inference efficiency. PAT reduces the computational cost of multi‑head attention and feed‑forward networks by using token aggregation‑expansion pairs: a token aggregation module (TAM) selects representative tokens based on 3D position anchors, and a token expansion module (TEM) propagates the learned representations back to the full token set. Combined with base‑sharing low‑rank adaptation (BSLoRA) for the TAMs, PAT achieves performance comparable to state‑of‑the‑art methods while requiring fewer trainable parameters and lower computational overhead.

By Zheng Liu, Xin Gao, Jinchao Zhu, Gao Huang
arXiv Computer Vision
3d ago

DiDA: Video Object Segmentation with Distillation Learning of Deformable Attention

DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.

By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung