Hugging Face Trending Papers

M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

Read the original on Hugging Face Trending Papers →

Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains constrained by high costs and privacy concerns, limiting its use in multimodal research and AI-driven applications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Aug 25

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

SynerMedGen is a unified framework that aligns medical multimodal understanding with generation tasks through task alignment. It introduces three generation‑aligned understanding tasks and a two‑stage training strategy that transfers representations learned during understanding to medical image synthesis. The model achieves strong zero‑shot performance on 22 synthesis tasks and outperforms state‑of‑the‑art specialized and unified models when combined with generation training, supported by a new 1M‑sample SynerMed dataset.

By Weiren Zhao, Yi Dong, Cheng Chen
arXiv Machine Learning
Sep 14

LatentVerse: A Framework for Understanding Shared and Modality-Specific Information in Multimodal Latent Representations

LatentVerse is a new framework that provides a web-based visual analytics platform and a command-line interface for analyzing multimodal latent representations. It unifies diagnostics for representation quality metrics and extends analysis to multimodal settings by decomposing embeddings into shared and modality-specific components. The authors evaluate the tool through simulations, real biomedical data analyses, and a user study, demonstrating its utility for interpretable evaluation of foundation model representations.

By Majd Alafrange, Samuel Friedman, John Kitonyo, Sana Tonekaboni, Mahnaz Maddah
arXiv Computer Vision
Sep 1

Coarse to Fine: Iterative Adversarial Neural Cellular Automata for Medical Image Synthesis

The paper introduces StyleGANCA, a lightweight neural cellular automata (NCA) based generative adversarial network designed for medical image synthesis. By combining a StyleGAN-inspired mapping network with adaptive style modulation in a multi-scale NCA framework, the model achieves high-quality image generation with far fewer parameters than existing adversarial, variational, diffusion, and NCA baselines. Experiments on BloodMNIST and PathMNIST show competitive FID and KID scores, and the synthetic images preserve class-specific information, effectively supporting downstream multi-class classifier training.

By Anh Thi Luu, Nick Lemke, Anirban Mukhopadhyay
arXiv Machine Learning
Sep 7

SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis

The paper introduces SMILE, a self‑explainable multimodal information bottleneck framework for medical diagnosis. It jointly optimizes predictive accuracy and modality‑specific explainability by selecting the most informative elements within each data modality. Experiments on diverse medical datasets show strong diagnostic performance, including a 9.1‑percentage‑point accuracy gain on the iCTCF dataset, and provide transparent, modality‑aware explanations that enhance both explainability and generalization.

By Yuqing Yang, Alexander Schmatz, Zhaozhao Ma, Changkyu Choi, Robert Jenssen, Shujian Yu