arXiv AI

OCT-FedSIR: Toward Trustworthy Federated Ophthalmic Learning under Annotation Noise

OCT-FedSIR is a reliability‑aware spectral framework designed for federated learning of OCT image classification in the presence of client‑dependent annotation noise and heterogeneous data distributions. It integrates class‑balanced spectral estimation, logit adjustment, complementary spectral descriptors, selective spectral relabeling, and noise‑aware federated optimization. Across 117 experimental conditions on three datasets, OCT‑FedSIR achieved a mean accuracy of 86.73%, outperforming RoFL (79.94%) and FedCorr (78.75%) and successfully identifying and correcting corrupted annotations with high precision.

arXiv AI
Sep 3

Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts

The paper evaluates federated learning with Low‑Rank Adaptation (LoRA) for fine‑tuning the BiomedCLIP vision‑language model on chest X‑ray classification across four international cohorts. Federated LoRA improves shared‑class AUC from 0.687 to 0.802, outperforming isolated single‑cohort training and approaching a centralized reference. The study shows that SVD‑based product‑space aggregation (FlexLoRA) is crucial for performance, while FedProx offers no advantage over FedAvg in this setting.

By Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal, Debesh Jha, Sunil Kumar Gaire
arXiv AI
Sep 11

OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

OmniMed‑FL is a multimodal federated learning framework that fuses chest radiographs and synthetic patient notes to classify five clinical conditions. The study benchmarks eight fusion strategies, three initializations, and four missing‑text imputation rules across 3–20 hospital clients under non‑IID Dirichlet partitioning, showing that federated approaches (FedAvg, FedProx, SCAFFOLD‑AdamW) outperform local‑only training. Multimodal fusion consistently improves performance, achieving macro‑F1 scores up to 0.956 on the synthetic corpus and 0.906 on the radiograph corpus.

By Ayush Debnath, Ruelia Saha, Sudip Misra
arXiv AI
Aug 11

Multimodal Federated Learning under Dual-Axis Modality Missingness

arXiv:2608. 09240v1 Announce Type: cross Abstract: Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual samples may contain only subsets of the modalities available locally.

By Adiba Orzikulova, Jaehyun Kwak, Jaemin Shin, Yunqi Guo, Xiaomin Ouyang, Guoliang Xing, Steven Euijong Whang, Sung-Ju Lee
arXiv Machine Learning
Aug 31

Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

The paper presents an explainable diabetic retinopathy classification framework that leverages vision foundation models—DINOv2, CLIP, and Vision Transformer—combined with various transfer learning techniques such as full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Using the ODIR dataset for internal validation and the APTOS dataset for external testing, DINOv2‑LoRA achieved the best internal AUROC (0.758) while DINOv2 and ViT full fine‑tuning reached the highest external AUROC (0.920). Explainability was assessed with Grad‑CAM and HiResCAM against expert‑annotated lesion masks from IDRiD, using Dice, IoU, and Pointing Game metrics, confirming that model attention aligns with clinically relevant retinal lesions.

By Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz
arXiv AI
Sep 1

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

Co-Annotator is a clinical AI system that distills expert gaze and dictation into two guidance components: a gaze‑aligned Vision Transformer that highlights fixation‑aligned areas of interest (AOIs) and an ontology‑bounded vision‑language model that pre‑fills editable biomarker summaries for retinal OCT. In controlled studies, each modality independently improved diagnostic accuracy and biomarker generation, and when combined across two academic institutions, the system increased correct diagnoses per minute by 40% and reduced comment editing time by 67% without compromising accuracy.

By Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor