The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.
By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
The paper tackles two main issues in multi-subject video generation—uncontrollable fidelity strength and semantic drift—by studying Diffusion Transformers (DiTs). It discovers that certain attention blocks naturally create an Intrinsic Spatial Grounding Map (ISGM) that accurately locates reference subjects. Leveraging this insight, the authors introduce Dual-phase Intrinsic Attention Leveraging (DIAL), which uses ISGM during low-noise stages to control fidelity strength without retraining and during high-noise stages to generate preference pairs for reinforcement learning, thereby anchoring attention and reducing semantic drift. Experiments on the OpenS2V-Eval benchmark show that DIAL outperforms baseline models, improving identity consistency and enabling controllable fidelity strength.
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing...
The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
By Jaymin Bhan, JiHong Jeon, SangYeop Jeong
arXiv:2608.30194v1 Announce Type: new
Abstract: Diffusion models have recently advanced text-to-video (T2V) generation, yet they still struggle with fine-grained compositional alignment, such as attr...
By Yujiang Pu, Yu Kong
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts.
The paper tackles two main issues in multi‑subject video generation—uncontrollable fidelity strength and semantic drift—by exploiting intrinsic attention patterns in Diffusion Transformers. It introduces an Intrinsic Spatial Grounding Map (ISGM) that accurately locates reference subjects and a Dual‑phase Intrinsic Attention Leveraging (DIAL) framework that uses ISGM during both training and inference. DIAL guides attention in low‑noise stages for precise fidelity control and builds preference pairs in high‑noise stages for reinforcement learning, resulting in superior identity consistency and controllable fidelity on the OpenS2V‑Eval benchmark.
By Niange Yu, Ye Tian, Biaolong Chen, Miao Lu, Aixi Zhang, Hao Jiang, Yunhai Tong, Pipei Huang
arXiv:2608. 03244v1 Announce Type: new Abstract: Image-goal visual navigation is a fundamental capability for embodied agents.
By Changqing Zhou, Yueru Luo, Zeyu Jiang, Changhao Chen
arXiv:2506. 00633v3 Announce Type: replace-cross Abstract: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space.
By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
CameraEditor is a new framework that transforms camera-controlled image editing into a temporal sequence prediction problem. By using video diffusion models, it incorporates a geometric perception module and dynamic reference routing to create precise visual references through dynamic panorama cropping. The method also inserts intermediate transition frames to handle large perspective shifts, maintaining content identity and spatial coherence, and is evaluated on a dataset of 5,760 instances with a benchmark of 462 test cases, achieving state‑of‑the‑art performance.
By Xin Shen, Chengyou Jia, Keshuo Xing, Zifeng Zhu, Changliang Xia, Bowen Ping, Zhuohang Dang, Hangwei Qian, Minnan Luo
The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.
By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv:2603. 28762v2 Announce Type: replace-cross Abstract: Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt.
By Omer Dahary, Benaya Koren, Daniel Garibi, Daniel Cohen-Or