arXiv Computer Vision

AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport

AlignMorph is a tuning‑free diffusion framework for image morphing that separates geometric alignment from generative denoising. It uses Global Semantic Transport—entropic optimal transport and reliability‑aware latent warping—to achieve diffusion‑compatible semantic alignment, and Coordinate‑Aligned Generation—symmetric bi‑phase attention handoff—to preserve spatial coordinates during denoising. The method eliminates ghosting and delivers superior structural coherence and temporal smoothness on morphing benchmarks without any per‑pair optimization.

arXiv AI
Jun 24

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

arXiv:2606. 24874v1 Announce Type: cross Abstract: Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks.

By Haorui Ji, Weizhe Liu, Hongdong Li, Hengkai Guo
arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv AI
Jun 16

Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment

arXiv:2606. 16799v1 Announce Type: cross Abstract: Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distortion dimensional conflict: monolithic representations optimized for semantic discrimination inherently entangle compositional understanding with low-level perceptual sensitivity, rendering them blind to fine-grained quality degradations.

By Zijie Meng
arXiv Computer Vision
Sep 15

Diffusion Trajectory Modeling for Semantic Correspondence

Diffusion Trajectory Modeling (DTM) treats the evolving feature maps of diffusion models as temporally structured trajectories rather than static snapshots. By interpreting each spatial patch’s progression across multiple timesteps as a trajectory, DTM captures semantic correspondence cues that prior methods miss. Experiments on SPair-71k, SPair-U, and AP-10K demonstrate that DTM achieves strong performance, highlighting the semantic value embedded in the diffusion process’s temporal axis.

By Yusung Choi
Hugging Face Trending Papers
Jun 23

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck.