Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv AI
3d ago

Transcriptome-informed multi-modal AI for predicting neoadjuvant therapy response from breast cancer biopsies

arXiv:2610.03693v1 Announce Type: new Abstract: Scarcity of labeled data limits development of deep learning biomarkers in oncology. We develop a two-stage AI model predicting pathological complete r...

By Jungkyu Park, Dhruva Biswas, Joseph Cappadona, Cerise Tang, Ken G. Zeng, Bartosz Machura, Chuwen Liu, Paolo Tarantino, Coral Omene, Francisco J. Esteva, Rohit Bhargava, Marcin Braun, Kamila Pa\'zdzierz, Jakub Czerwi\'nski, Hanna Roma\'nska-Knight, Albert Grinshpun, Bareket Daniel, Michele Buchinger, Frederick Howard, Piotr Wysocki, Brie Chun, Freya Schnabel, Rich Caruana, Jan Witowski, Krzysztof J. Geras
arXiv Machine Learning
5d ago

MOVE: Multimodal Open-world Verification and Expansion for Graph Learning

The paper introduces MOVE, a framework for multimodal open‑world verification and expansion in graph learning. MOVE jointly uses visual tokens, textual attributes, and graph context to identify nodes that cannot be assigned to existing classes, then employs a multimodal LLM to generate candidate class descriptions. It selectively expands the class space only when multimodal evidence consistently supports the new classes, avoiding redundancy, and reports an average 11.87% improvement across unknown recognition, open‑domain annotation, and downstream graph learning tasks.

By Zekai Chen, Jiayang Xing, Xun Wu, Miao Zhang, Xunkai Li, Kairui Yang, Zhengyu Wu, Xu Wang, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
5d ago

Beyond Unimodal Bases: Pullback Geometry for Multimodal Data

The paper introduces a pullback Riemannian geometry tailored for multimodal data by employing a latent Gaussian mixture model. It defines a smooth, positive‑definite metric based on responsibility‑weighted component precision, extending the standard single‑Gaussian construction. Experiments on synthetic, multi‑view image, and MNIST datasets demonstrate reduced transport distortion, accurate trajectory recovery, and more realistic interpolation.

By Honglei Brinkmann, Lucas Ng, Georgios Batzolis, Mark Girolami, Carola-Bibiane Sch\"onlieb, Willem Diepeveen
arXiv Machine Learning
5d ago

HADRec: A Hierarchy-Aware Drug Recommendation Framework by Fusing Molecular Knowledge and Electronic Health Record

HADRec is a Hierarchy-Aware Drug Recommendation framework that fuses molecular knowledge and electronic health records to improve medication recommendation. It uses LLaMA-7B to encode clinical notes, ChemBERTa to encode drug SMILES strings, and a cross‑attention mechanism for multimodal fusion, while a hierarchical predictor and consistency constraint loss enforce adherence to the ATC classification system. Experiments on MIMIC‑III and MIMIC‑IV show state‑of‑the‑art performance, strong generalization, and well‑calibrated predictions, with counterfactual evaluation indicating clinically aligned reasoning.

By Junke Wang, Hongshun Ling, Li Zhang, Jinjing Wu, Tong Shao, Fang Wang, Yuan Gao
arXiv Machine Learning
5d ago

Debias Anything: Fairness with Diversity without Supervision in Diffusion Models

The paper introduces a method called Debias Anything that jointly addresses fairness and diversity in diffusion models without requiring sensitive-attribute annotations. By connecting a frozen diffusion model to a pretrained vision-language embedding space via an adapter, the approach uses pairs of text prompts to guide batch composition toward desired attribute proportions and employs a disagreement score to promote diversity. The method is applicable to both unconditional and text-conditional diffusion models and demonstrates improved quality and diversity while maintaining comparable fairness levels in experiments.

By Th\'eau d'Audiffret, Mariia Vladimirova, Jean-Yves Franceschi