arXiv:2607. 04079v1 Announce Type: cross Abstract: Recent Multi-modal Large Language Models (MLLMs) have demonstrated remarkable performance on 2D question answering tasks.
By Ruei-Chi Lai, Bolivar Solarte, Chin-Hsuan Wu, Yi-Hsuan Tsai, Min Sun
arXiv:2609.16233v1 Announce Type: cross
Abstract: Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current ben...
By Anubhav Khanal, Prabigya Acharya, Roshni Poudel, Sujan Kapali, Bigyan Bhatta, Pramish Paudel, Francois Rameau, Danda Pani Paudel
arXiv:2609.15137v1 Announce Type: cross
Abstract: 3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic f...
By Davit Soselia, Joseph JaJa, Amitabh Varshney
arXiv:2608. 01185v1 Announce Type: cross Abstract: Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial reasoning for tasks such as 3D question answering.
By Changwoo Baek, Kyeongbo Kong
arXiv:2610.00040v1 Announce Type: new
Abstract: Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation mo...
By Thanh-Khoi Nguyen, Thien-Phuc Tran, Minh-Triet Tran
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clo...