arXiv:2608. 00111v1 Announce Type: cross Abstract: Image restoration quality can be evaluated along three complementary facets: pixel-level fidelity, human perception, and downstream machine preference.
By Kuan-Yen Chen, Fang-Yi Su, Philip Chikontwe, Jung-Hsien Chiang
arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.
By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo
arXiv:2604. 05182v2 Announce Type: replace-cross Abstract: We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.
By Zhengqin Li, Cheng Zhang, Jakob Engel, Zhao Dong
arXiv:2609.00663v1 Announce Type: new
Abstract: Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separatin...
By Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban
arXiv:2608.23549v1 Announce Type: new
Abstract: Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces...
By Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan
arXiv:2509. 07295v4 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture.
By Ji Xie, Trevor Darrell, Luke Zettlemoyer, XuDong Wang