arXiv Computer Vision

PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization

arXiv Computer Vision
Sep 7

Compositional Reward Models for Conditional Medical Image Generation

The paper introduces PRISM, a Compositional Reward Model framework that decomposes image quality into multiple verifier‑grounded stages for conditional medical image generation. By assigning distinct rewards for fine‑to‑coarse properties—such as intensity, texture, structural alignment, and semantic fidelity—and combining them via a Hierarchical Constrained Propagation mechanism, PRISM addresses shortcomings of single‑scalar reward approaches. Experiments on PanNuke, CeDeM, and ISIC datasets show that data generated with PRISM improves downstream model performance, achieving higher mDice, lower MRE, and increased F1 scores compared to baseline methods.

By Aayush Kumar Tyagi, Prathosh A. P., Mausam
arXiv Machine Learning
Aug 20

Simple, Safe, and Overlooked: Reclaiming Sustainable Domain Generalization with Statistical Color Matching

The paper introduces Colorist, a data‑augmentation method that uses classical statistical color matching to generate domain‑shifted medical images. By applying global mean‑standard‑deviation matching in RGB space, Colorist creates structurally intact variations without neural networks, outperforming deep generative models in fidelity and color alignment. Across multiple medical imaging datasets, it boosts balanced accuracy by up to 9% over state‑of‑the‑art domain‑generalization regularizers and 13% over no augmentation, while reducing computational cost and preserving anatomical structure.

By Sebastian Doerrich, Francesco Di Salvo, Shyam Nandan Rai, Marco Lents, Christian Ledig
arXiv AI
Aug 20

OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation

OptiModNet is a lightweight UNet‑Transformer hybrid designed for optic disc and cup segmentation. It incorporates grouped‑query and channel attention across multiple stages, along with an Aggregated Pyramid Loss to improve gradient flow and structural consistency. Evaluated on the REFUGE2 dataset, it surpasses existing methods by over 2.5 % while using only 3.73 GFLOPs and 1.93 M parameters.

By Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
arXiv Machine Learning
Aug 5

Contrast-invariant deep ptychography neural networks

arXiv:2608. 02869v1 Announce Type: new Abstract: Ptychography neural networks suffer from scaling inconsistencies when generalizing out of distribution, limiting their real world viability.

By Albert Vong, Steven Henke, Oliver Hoidn, Hanna Ruth, Junjing Deng, Apurva Mehta, David Shapiro, Alexander Hexemer, Nicholas Schwarz
arXiv Computer Vision
5d ago

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.

By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv Computer Vision
Aug 27

Improving Cross-Site Whole-Heart Segmentation

The paper presents a modality‑routed 3D cardiac segmentation pipeline that combines TotalSegmentator‑initialized nnU‑Netv2 models with site‑characterized, label‑preserving appearance augmentation. By analyzing measurable image properties across sites, the authors design a bias‑field plus Bezier augmentation strategy that smooths spatial intensity perturbations and remaps intensities nonlinearly, followed by class‑wise largest‑connected‑component cleanup. On held‑out validation splits, this approach raises CT mean Dice from 0.8350 to 0.9135 and MRI mean Dice from 0.7695 to 0.7830 while reducing HD95, demonstrating improved cross‑site robustness in limited‑data whole‑heart segmentation.

By Tanish Mudaliar, Justin Li, Daniel Lin, Julianna Vo, Kaitao Liao, Xin Wang, Shu Hu