arXiv:2608. 07579v1 Announce Type: cross Abstract: The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only.
By Abdullah Naeem, Anav Katwal, Ayon Dey, Noman Khan, Md Tamjidul Hoque
A surveillance camera is an image sensor whose silent physical degradation invalidates every downstream consumer of its data. In-situ integrity alarms for such vision sensors require low false-alarm rates, bounded computation, and diagnosable behavior under nuisance illumination changes.
arXiv:2608.29705v1 Announce Type: cross
Abstract: Feed-forward 3D reconstruction models emit a per-pixel confidence that downstream systems read as a reliability signal. It is trained as a loss weigh...
By Nanxing Nick Deng, Qing Cheng, Niclas Zeller, Daniel Cremers
arXiv:2606. 16479v1 Announce Type: cross Abstract: Visual Geometry Grounded Transformer (VGGT) has already attracted a great deal of attention in a short period of time, not least due to the Best Paper Award at CVPR-2025.
By Markus Hillemann, Robert Langend\"orfer, Steven Landgraf, Markus Ulrich
arXiv:2607. 22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting.
By M M Asif Ferdous
arXiv:2608. 05670v1 Announce Type: new Abstract: A model's agreement across perturbed inputs is used both as a label-free reliability signal and as a self-training target, on the premise that agreement tracks correctness.
By Rasul Khanbayov, Hasan Kurban
arXiv:2606. 18451v1 Announce Type: new Abstract: Single-image-to-3D generators are improving quickly, but there is no agreed, human-free way to tell whether one generated mesh is better than another.
By Ali Asaria, Tony Salomone, Deep Gandhi
arXiv:2608.21402v1 Announce Type: cross
Abstract: World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera...
By Bingqi Huang, Bingchuan Wei, Yingkai Cai, Zhaokui Wang
arXiv:2608.29193v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. We propose HalluPrism, a behavioral...
By Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty
arXiv:2608.22300v1 Announce Type: new
Abstract: Co-registration underlies nearly every multi-temporal and multi-sensor use of optical satellite imagery, and operational products still carry documente...
By Shoukun Sun, Zhe Wang, Sanaz Salati, Jiyin Zhang, Hui Wang, Xiaogang Ma
arXiv:2608.29680v1 Announce Type: new
Abstract: Feed-forward 3D foundation models reconstruct perspective scenes in one pass. Satellite photogrammetry needs a different product, one that domain adapt...
By Zhe Dong, Wanqing Wu, Yuzhe Sun, Haochen Jiang, Yuchen Ma, Lecheng Ren, Tianzhu Liu, Yanfeng Gu
SNF-Bench is an evaluation framework for long‑horizon fixed‑camera video generation that separates static background fidelity from dynamic flow persistence and drift leakage. It reports these three factors independently, using controlled injections of translation, rotation, scale drift, and progressive freezing to validate each metric’s sensitivity. Auditing public checkpoints shows that whole‑frame motion metrics can mislead, while SNF‑Bench reveals the true trade‑offs between motion quality and background stability.
By Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park