Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision
arXiv:2505. 12532v3 Announce Type: replace-cross Abstract: Efficiently adapting large pretrained models is critical under tight compute and memory budgets.
arXiv:2505. 12532v3 Announce Type: replace-cross Abstract: Efficiently adapting large pretrained models is critical under tight compute and memory budgets.
arXiv:2505. 18315v3 Announce Type: replace-cross Abstract: We introduce \textbf{CoLoRA} (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural networks (CNNs).
arXiv:2608. 07749v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace.
arXiv:2609.14437v1 Announce Type: cross Abstract: Deepfake detection systems often exhibit significant performance degradation when deployed on unseen manipulation methods, limiting their reliability...
arXiv:2608.21300v1 Announce Type: new Abstract: Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot se...
The paper introduces an Early Intervention (EI) framework for multimodal medical image classification that addresses two key challenges: limited exploitation of complementary multimodal information and scarcity of labeled data for Vision Foundation Models (VFMs). EI treats one modality as the target and uses high‑level semantic tokens from other modalities as intervention tokens to guide the target’s embedding early in the process. The authors also propose Mixture of varied‑rank LoRAs (MoR) for efficient VFM adaptation, and demonstrate the method’s effectiveness on retinal, skin, and knee medical image datasets.
arXiv:2608.22619v1 Announce Type: cross Abstract: Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-t...
The paper evaluates federated learning with Low‑Rank Adaptation (LoRA) for fine‑tuning the BiomedCLIP vision‑language model on chest X‑ray classification across four international cohorts. Federated LoRA improves shared‑class AUC from 0.687 to 0.802, outperforming isolated single‑cohort training and approaching a centralized reference. The study shows that SVD‑based product‑space aggregation (FlexLoRA) is crucial for performance, while FedProx offers no advantage over FedAvg in this setting.
arXiv:2608.21571v1 Announce Type: new Abstract: Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage...
Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation (MRSeg) is a parameter‑efficient framework that uses frozen ConvNeXt‑Tiny and PubMedBERT encoders to extract multiscale visual features and clinical text tokens. A joint router predicts a sparse mixture over low‑rank adapter bases, enabling separate adaptation for two visual scales and text while keeping feature‑specific parameters distinct. Region Bridge aggregates dense visual tokens into latent regions using text‑derived queries, refines them via self‑attention and text cross‑attention, and redistributes the refined information back to the feature maps, culminating in a multiscale decoder that combines refined semantic features with shallow image evidence. MRSeg achieves state‑of‑the‑art Dice/mIoU scores on QaTa‑COV19 and MosMedData+ with only 7.11 M trainable parameters and 7.60 GFLOPs.
arXiv:2412. 10362v2 Announce Type: replace Abstract: Low-rank adapters (LoRA) enable finetuning of large models with only a small number of parameters.
Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language fe...