arXiv Machine Learning By Bowen Wang, Youwen Zhang, Ritesh Mehta

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

Read the original on arXiv Machine Learning →

arXiv:2607. 27763v1 Announce Type: cross Abstract: We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 30

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.

arXiv Computer Vision
Sep 15

Designing UNICORN: a Unified Benchmark for Imaging in Computational Pathology, Radiology, and Natural Language

arXiv:2603.02790v2 Announce Type: replace Abstract: Foundation models are changing the way we develop medical artificial intelligence. By learning broadly generalizable features across diverse data m...

By Michelle Stegeman (and on behalf of the UNICORN consortium), Lena Philipp (and on behalf of the UNICORN consortium), Fennie van der Graaf (and on behalf of the UNICORN consortium), Marina D'Amato (and on behalf of the UNICORN consortium), Cl\'ement Grisi (and on behalf of the UNICORN consortium), Luc Builtjes (and on behalf of the UNICORN consortium), Joeran S. Bosma (and on behalf of the UNICORN consortium), Judith Lefkes (and on behalf of the UNICORN consortium), Rianne A. Weber (and on behalf of the UNICORN consortium), James A. Meakin (and on behalf of the UNICORN consortium), Thomas Koopman (and on behalf of the UNICORN consortium), Anne Mickan (and on behalf of the UNICORN consortium), Mathias Prokop (and on behalf of the UNICORN consortium), Ewoud J. Smit (and on behalf of the UNICORN consortium), Fr\'ed\'erique Meeuwsen (and on behalf of the UNICORN consortium), Geert Litjens (and on behalf of the UNICORN consortium), Jeroen van der Laak (and on behalf of the UNICORN consortium), Bram van Ginneken (and on behalf of the UNICORN consortium), Maarten de Rooij (and on behalf of the UNICORN consortium), Henkjan Huisman (and on behalf of the UNICORN consortium), Colin Jacobs (and on behalf of the UNICORN consortium), Francesco Ciompi (and on behalf of the UNICORN consortium), Alessa Hering (and on behalf of the UNICORN consortium)
arXiv Computer Vision
Aug 27

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

The paper introduces a label‑free method called AURCC for selecting the best foundational model for medical image classification when the target domain lacks labels. AURCC uses a pseudo‑label discrepancy computed by the SUDO framework to score models without fine‑tuning. Experiments on chest X‑ray data across three inter‑hospital shifts show that AURCC closely matches the true model ranking, outperforming simple source‑accuracy baselines especially when source data are limited.

By Juan I\~naki Larrea, Lucas Mansilla, Enzo Ferrante
arXiv AI
Aug 26

OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses

The paper introduces D3-Omni, a balanced and decoupled benchmark designed to diagnose fine‑grained multimodal understanding in OmniJudges that evaluate text‑to‑image, text‑to‑video, and text‑to‑speech generation. D3-Omni covers 53 orthogonal binary dimensions across 10,671 samples, using fixed positive seeds and controlled prompt rewriting to generate negatives, thereby ensuring each error can be attributed to a single capability. The benchmark’s dual‑balanced, decoupled, and dynamic design achieves near 1:1 per‑dimension parity and a uniform total‑score distribution, revealing that strong OmniJudges often miss modality‑related failures and treat distinct attributes as a single decision, masking systematic blind spots.

By Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang, Ziyi Cheng, Xinfa Zhu, HangRui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu