arXiv:2608.28707v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with quest...
By Anoop Senthil
arXiv:2606. 04772v1 Announce Type: cross Abstract: Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience.
By Hoang-Son Vo, Van-Hung Bui, Minh-Huy Mai-Duc, Tien-Dung Mai, Soo-Hyung Kim
arXiv:2606. 00275v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have demonstrated impressive performance on multimodal tasks through scaled architectures and extensive training.
By Zijie Zhou, Dandan Zhu, Hangxiangpan Wang, Heng Zhang, Huishen Jiao, Yi Zhao
arXiv:2608. 10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios.
By Carlos Zamora, Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
The paper introduces TEAMS, a vision‑language Mamba snake framework that enhances deep snake instance segmentation. It adds a Spatiotemporal Snake Evolution Strategy to handle complex shapes, a Contour Morphology‑Aware Mamba to improve fine‑grained detail capture, and a Text‑prompted Collaborative Dual‑Head Snake to integrate textual cues and reduce detection errors. Experiments on five medical imaging datasets show TEAMS surpasses existing methods, achieving significant gains in mDice and mBF metrics.