arXiv Machine Learning

LatentVerse: A Framework for Understanding Shared and Modality-Specific Information in Multimodal Latent Representations

LatentVerse is a new framework that provides a web-based visual analytics platform and a command-line interface for analyzing multimodal latent representations. It unifies diagnostics for representation quality metrics and extends analysis to multimodal settings by decomposing embeddings into shared and modality-specific components. The authors evaluate the tool through simulations, real biomedical data analyses, and a user study, demonstrating its utility for interpretable evaluation of foundation model representations.

arXiv AI
Aug 26

MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models

MolEmb is a lightweight framework that adapts multimodal large language models (MLLMs) to serve as general molecular embedding models. By aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, MolEmb produces embeddings conditioned on both a molecular profile and a natural‑language semantic context. The model performs competitively on molecular property prediction and enables cross‑modal molecule‑text retrieval, while the newly introduced MolCAR benchmark demonstrates that context‑aware molecular embedding is largely a data property of the supervision.

By Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu
Hugging Face Trending Papers
Jul 23

M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains constrained by high costs and privacy concerns, limiting its use in multimodal research and AI-driven applications.

arXiv Machine Learning
Jul 9

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

arXiv:2607. 07673v1 Announce Type: cross Abstract: Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams.

By Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
arXiv Computation and Language
Aug 28

Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval

Case2Flow is a new task that retrieves the most relevant guideline flowchart for a given patient case from a collection of medical guideline documents. The authors created FlowAtlas, a curated corpus of 202 flowcharts extracted from 2,080 guidelines, and a pipeline that generates 1,911 aligned case‑flowchart pairs. Their evaluation shows that existing multimodal retrieval methods often overrely on keywords and spurious token‑patch matches, and they propose CRISP, a training‑free scoring method that improves Recall@1 by up to 18.71 percentage points and gains preliminary feasibility evidence from a blinded physician assessment.

By Jiale Wei, Yufan Chen, Alexander Jaus, Zdravko Marinov, Julian Friedrich, Simon Rei{\ss}, Jens Kleesiek, Rainer Stiefelhagen
arXiv Machine Learning
Jul 27

LatentFlow: Visual Analytics for Latent Space Analysis in Molecular Graph Neural Networks

arXiv:2607. 21941v1 Announce Type: new Abstract: Chemists and materials scientists increasingly use machine learning models, such as graph neural networks (GNNs), to predict properties of molecules and the outcomes of their reactions.

By Shiyi Liu, Jiaqing Chen, Nicholas Hadler, Rostyslav Hnatyshyn, Michael W. Mahoney, Talita Perciano, John F. Hartwig, Gunther H. Weber, Ross Maciejewski
arXiv Machine Learning
Aug 6

Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals

arXiv:2510. 09764v2 Announce Type: replace Abstract: Modeling multi-modal time-series data is critical for capturing system-level dynamics, particularly in biosignals where modalities such as ECG, PPG, EDA, and accelerometry provide complementary perspectives on interconnected physiological processes.

By Wanting Mao, Maxwell A Xu, Harish Haresamudram, Mithun Saha, Santosh Kumar, James Matthew Rehg