Two-Stage Mixture-of-LoRA for Multi-Task Medical Vision-Language Learning
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper evaluates federated learning with Low‑Rank Adaptation (LoRA) for fine‑tuning the BiomedCLIP vision‑language model on chest X‑ray classification across four international cohorts. Federated LoRA improves shared‑class AUC from 0.687 to 0.802, outperforming isolated single‑cohort training and approaching a centralized reference. The study shows that SVD‑based product‑space aggregation (FlexLoRA) is crucial for performance, while FedProx offers no advantage over FedAvg in this setting.
arXiv:2607. 20641v1 Announce Type: new Abstract: Federated learning (FL) enables multiple clinical institutions to collaboratively train a shared disease classifier without centralizing patient data.
arXiv:2607. 11839v1 Announce Type: cross Abstract: This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments.
Vision-language models (VLMs) such as CLIP enable zero-shot classification by comparing image features with text prompts in a shared embedding space. A fundamental property underlying this capability is the global comparability of logits across arbitrary candidate classes.
This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated in stages without retraining previously learned components.
arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.