arXiv AI

UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning

UniAR is a unified framework that improves autism spectrum disorder (ASD) recognition by using multi-granularity prompt learning and a large multimodal model to generate diagnostic descriptions at word, phrase, and sentence levels. It aligns these semantic representations with visual evidence through a Mixture-of-Experts-based Multi-Scale Alignment Module, enabling robust ASD detection across heterogeneous data types. Experiments on four brain MRI and facial expression benchmarks show that UniAR outperforms state‑of‑the‑art methods, achieving 75.9% accuracy on MRI and 91.6% on facial benchmarks, with gains of 1.5 and 1.2 percentage points respectively.

arXiv AI
Aug 20

A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities

The article reviews 55 machine‑learning studies on autism spectrum disorder (ASD) published between 2017 and 2023, focusing on how ML can aid early diagnosis and treatment. It finds that supervised learning dominates current research, while deep learning is gaining traction as data volumes grow. The review highlights the need for models that fuse complex data—such as genetic, clinical, wearable, and biometric sources—to improve diagnostic accuracy and enable continuous, non‑intrusive monitoring.

By Rafael Mu\~noz-Terol, Jes\'us Peral, Sandra Amador, David Gil
arXiv Machine Learning
Jul 17

Multimodal Semantic-Aware Contrastive Learning For False Negative Mitigation in 3D Medical Imaging

arXiv:2607. 14995v1 Announce Type: new Abstract: Multimodal Contrastive Learning (CL) has shown significant performance in aligning representations across various data modalities and improving downstream tasks, especially in healthcare.

By Sara Ketabi, Matthias W. Wagner, Cynthia Hawkins, Uri Tabori, Birgit Betina Ertl-Wagner, Farzad Khalvati
arXiv AI
Aug 14

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

arXiv:2608. 12689v1 Announce Type: cross Abstract: Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically.

By Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
arXiv AI
Sep 1

AOI-Net: Structural Face AOI-Guided Eye-Gaze Track Representation Learning for Autism Spectrum Disorder Detection

AOI-Net introduces a structural face AOI-guided Eye‑Gaze Track Network that jointly models short‑term temporal dynamics and AOI‑level structural organization for Autism Spectrum Disorder detection. The network uses a gating mechanism to adaptively combine complementary representations and incorporates class‑distribution‑aware learning to address the imbalance between ASD and typically developing participants. Experiments on a large clinical eye‑tracking database with over 1,300 participants demonstrate that AOI‑Net outperforms state‑of‑the‑art methods and offers interpretable gaze‑behavior modeling for scalable AI‑driven ASD screening.

By Zhanpei Huang, Binbin Sun, Jialiang Chen, Yiou Wang, Taochen Chen, Yuzhu Ji, Yiqun Zhang, Yiu-Ming Cheung
arXiv AI
Aug 26

Screening Autism Spectrum Disorder in children using Deep Learning Approach : Evaluating the classification model of YOLOv26s by comparing with other models

arXiv:2306.14300v2 Announce Type: replace-cross Abstract: Autism spectrum disorder (ASD) is a developmental condition that presents significant challenges in social interac- tion, communication, and...

By Subash Gautam, Sagar Pathak, Prabin Sharma, Bidhya Shrestha, Kisan Thapa, Shubham Joshi, Mala Deep Upadhaya, Dikshya Thapa, Chandiprasad Chintalapati, Sagar Duwal, Angela Upreti, Salik Ram Khanal
arXiv AI
Jul 7

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation

arXiv:2607. 04344v1 Announce Type: cross Abstract: While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like ocular surface diseases (OSDs) is severely hindered by a paucity of high-fidelity, multimodal instruction-tuning data.

By Hao Wei, Wenjin Qi, Dasen Dai, Minqing Zhang, Wu Yuan
arXiv Computation and Language
Sep 22

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

Lingshu is a medical‑specialized multimodal large language model that addresses key limitations of existing medical MLLMs, such as narrow knowledge coverage, hallucinations, and weak reasoning. The authors curate a comprehensive dataset combining medical imaging, texts, and general‑domain data, then train Lingshu in multiple stages to embed medical expertise and improve task performance. They also introduce MedEvalKit, a unified evaluation framework, and demonstrate that Lingshu outperforms current open‑source multimodal models on multimodal QA, text‑based QA, and medical report generation.

By Weiwen Xu, Hou Pong Chan, Long Li, Mahani Aljunied, Ruifeng Yuan, Jianyu Wang, Chenghao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, Yu Sun, Junao Shen, Chaojun Wang, Jie Tan, Deli Zhao, Tingyang Xu, Hao Zhang, Yu Rong
arXiv Machine Learning
Sep 7

SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis

The paper introduces SMILE, a self‑explainable multimodal information bottleneck framework for medical diagnosis. It jointly optimizes predictive accuracy and modality‑specific explainability by selecting the most informative elements within each data modality. Experiments on diverse medical datasets show strong diagnostic performance, including a 9.1‑percentage‑point accuracy gain on the iCTCF dataset, and provide transparent, modality‑aware explanations that enhance both explainability and generalization.

By Yuqing Yang, Alexander Schmatz, Zhaozhao Ma, Changkyu Choi, Robert Jenssen, Shujian Yu
arXiv AI
Aug 28

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.

By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao