arXiv Computer Vision By Zhiwen Yang, Jiayin Li, Chengyu Liu, Hui Zhang, Bingzheng Wei, Yan Xu

UniH$^3$: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration

Read the original on arXiv Computer Vision →

UniH$^3$ is a new framework for all-in-one medical image restoration that unifies hierarchical homogeneity and heterogeneity. It introduces a Hierarchical Homogeneity Memory (H2M) module to distill and retrieve shared anatomical priors, and a Hierarchical Heterogeneity Balancer (H2B) to mitigate inter- and intra-task conflicts during training. Experiments on MedIR-2D-500K and MedIR-3D-3K show that UniH$^3$ achieves state‑of‑the‑art performance for both multi‑task and single‑task restoration.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 25

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

SynerMedGen is a unified framework that aligns medical multimodal understanding with generation tasks through task alignment. It introduces three generation‑aligned understanding tasks and a two‑stage training strategy that transfers representations learned during understanding to medical image synthesis. The model achieves strong zero‑shot performance on 22 synthesis tasks and outperforms state‑of‑the‑art specialized and unified models when combined with generation training, supported by a new 1M‑sample SynerMed dataset.

By Weiren Zhao, Yi Dong, Cheng Chen
arXiv AI
Aug 11

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.

By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
arXiv Computer Vision
Sep 11

MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation

MedGEN-Bench is a new benchmark for open‑ended multimodal medical generation that addresses limitations in current medical visual benchmarks, such as query‑image misalignment, closed‑ended answer spaces, and text‑centric outputs. The dataset contains 6,422 image‑text pairs across six imaging modalities, 15 clinical tasks, and 27 subtasks, including VQA, image editing, and contextual multimodal generation pairs. Evaluation combines reference‑based fidelity metrics with a structured, checklist‑guided assessment by a medical VLM judge, and preliminary results show that image‑output tasks remain unsaturated while contextual augmentation improves image‑instruction similarity.

By Junjie Yang, Yuhao Yan, Gang Wu, Rui Qian, Zhisheng Chen, Haijiang Li, Yuhe Wu, Qichao Zhao, Dawen Tian, Xiang Wan, Fenglei Fan, Wenjian Qin, Yongquan Zhang, Feiwei Qin, Changmiao Wang