Hugging Face Trending Papers

Multilingual Hematology Visual Question Answering Dataset

Read the original on Hugging Face Trending Papers →

Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such as Visual Question Answering. However, existing hematology vision-language resources remain predominantly English centric, limiting their applicability in multilingual healthcare environments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
1d ago

PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology

PathLang is a new, language‑centered benchmark for vision‑language models in computational pathology. It keeps the underlying pathology slides, ground‑truth labels, and image‑text evaluation direction constant while systematically varying the diagnostic language used in prompts, reflecting how pathologists actually rephrase diagnoses. The benchmark covers four task families—zero‑shot classification with image‑text alignment, cross‑modal retrieval, paraphrase robustness, and open‑vocabulary diagnosis retrieval—and evaluates nine VLMs across five public datasets spanning four organs.

By Fanqi Cheng, Kuo Gong, Shangke Liu, Beidi Zhao, Junchao Zhu, Zheyu Zhu, Leiyue Zhao, Fengbei Liu, John Cannon, Gang Wang, Zu-hua Gao, Kenji Ikemura, Yihe Yang, Yaohong Wang, Yuankai Huo, Xiaoxiao Li, Mert R. Sabuncu, Ruining Deng
arXiv Machine Learning
Jun 5

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

arXiv:2606. 05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.

By Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Mohamed Hefeeda, Ehsaneddin Asgari
arXiv AI
Jun 12

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

arXiv:2606. 13572v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios.

By Tanmoy Kanti Halder, Akash Ghosh, Subhadip Baidya, Arijit Roy, Sriparna Saha