arXiv:2606. 20364v1 Announce Type: new Abstract: A companion study established a de-biased, cross-model VLM-as-3D-judge that reliably ranks single-image-to-3D mesh quality where cheap geometry and CLIP proxies fall short.
By Ali Asaria, Tony Salomone, Deep Gandhi
VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.
By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automated judge depends on the entire evaluation pipeline, not only the underlying vision-language model (VLM), but also how assets are rendered, what visual evidence is provided, how the task is specified, and how human reference labels are constructed.
arXiv:2607. 10826v1 Announce Type: cross Abstract: Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow.
By Zhenyu Zhao, Nanshan Jia, Jihyeon Je, Yifu Tang, Alvin Chan, Michael Spedden, Michael V. Palleschi, Sui Huang, Jingshen Wang, Zeyu Zheng
arXiv:2605.31351v2 Announce Type: replace-cross
Abstract: AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradig...
By Yi Zhao, Siqi Wang, Zhe Hu, Yushi Li, Jing Li
arXiv:2606.29167v2 Announce Type: replace
Abstract: Dense correspondence on in-the-wild 3D scans must handle severe non-isometric deformation, partial observations, topology artifacts, irregular disc...
By Qilong Liu, Qinfeng Xiao, Chenyuan Yi, Yongsheng Lin, Liying Zhang, Kit-lun Yick
The paper introduces a dual‑judge evaluation protocol for vision‑language models in legally grounded tasks, pairing a 0‑10 quality judge with a strict binary semantic‑equivalence judge. Using a controlled UK traffic‑sign interpretation task, the authors analyze 4,680 evaluations across visibility and occlusion conditions, finding moderate association between judges and an asymmetric Type II error pattern that is most pronounced under heavy occlusion. The protocol requires only one additional LLM call and reveals quality‑trustworthiness signals that single‑judge methods miss.
By Su Myat Noe, Ha Thanh Nguyen, May Myo Zin, Ken Satoh
The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.
By Amit Jadhav, Shaurya Beriwala, Beomjin Kim
arXiv:2609.26168v1 Announce Type: cross
Abstract: Recent work reports that vision--language models (VLMs) struggle to establish and maintain stable reference in repeated reference games. Rather than...
By Joseph Bingham
arXiv:2606. 19162v1 Announce Type: new Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself.
By Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana Romero-Soriano, Michal Drozdzal
ReVisIT is a train‑free framework that turns retrieved image‑label pairs into units of visual thought, combining structured class definitions, multimodal retrieval, and alternating user/assistant injection before joint decoding. On several benchmarks—including Fast Open MiniImageNet, Bongard‑OpenWorld, and the newly released MAAC‑Bench—ReVisIT achieves performance comparable to or surpassing large, trained models while using far fewer parameters. The approach demonstrates that high‑quality retrieval and a simple turns layer can provide a universal performance boost across diverse multimodal tasks.
By Bingchen Huang, Zhiling Wang, Yifu Chen, Yuanchao Du
The paper introduces SphereTrust, a method that uses frozen self‑supervised hyperspherical features to evaluate and rank candidate masks produced by foundation segmenters like SAM. By measuring angular contrast, foreground coverage, and image‑frame contact, SphereTrust can select high‑quality masks in 0.55 s per image and outperforms existing baselines on multiple segmentation tasks. The selected masks are then used as priors to train student models, improving performance on several benchmark datasets.
By Xinge Guo, Fengyang Xiao, Dingming Zhang, Yuhan Chen, Rihan Zhang, Xingjian Li, Tianyang Wang, Chunming He, Sina Farsiu