3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency.
arXiv:2609.38689v1 Announce Type: new
Abstract: Immersive displays can enable rich and diverse virtual experiences. Manually authoring every possible experience to realize this potential, however, is...
By Debabrata Mandal, Dongdong Fu, Jonathon Miller, William Villareal, Xi Peng, Praneeth Chakravarthula
The paper introduces a 2D Gaussian Splatting pipeline that renders a dominant-eye RGB image and depth proxy, then reprojects and selectively patches the affiliated eye to reduce redundant work. By reusing alpha-blending weights and generating adaptive regions of interest, the method cuts sequential binocular rendering time by 15.5% to 28.8% and GPU memory by 6% to 11% on several datasets, with minimal quality loss. It demonstrates a practical efficiency‑quality trade‑off for static‑scene stereo rendering and suggests further evaluation on dynamic scenes and VR hardware.
By Hongfei Zhu, Ling Zhou
The paper introduces a multidimensional observer model that represents images as distributions in a latent perceptual space and models human image quality judgment as comparisons of noisy samples. By aligning the model with neural representations in the primate ventral stream and fitting it to large-scale behavioral data, the authors demonstrate that the perceptual space required for human quality assessment is extremely low-dimensional relative to the image space. The study reveals that the structure of this perceptual space differs between low-level and high-level quality judgments, indicating that humans construct task-dependent perceptual spaces during visual decision making.
By Sheng Zhao, Weikai Lin, Yuhao Zhu
Low-light image enhancement algorithms (LIEAs) aim to improve the visibility of images captured under poor illumination. However, the enhancement process often introduces artifacts such as noise amplification, color shift, structural damage, and over-exposure, which degrade the perceptual quality of the enhanced images.
arXiv:2609.34367v2 Announce Type: replace
Abstract: Learned image codecs (LICs) achieve high reconstruction quality, but their decoding speed is often insufficient for immersive virtual reality (VR)....
By Yulong Cheng, Youneng Bao, Junfeng Zhou, Mu Li, Jie Wen
The paper introduces PIMDE, a self‑supervised monocular depth estimation framework that decomposes input images into perceptual feature maps, each encoding a specific visual cue. Separate depth branches process these maps to produce individual depth estimates, which are then fused explicitly. Experiments on the KITTI benchmark show that PIMDE matches the accuracy of existing self‑supervised methods while offering clearer insight into how each perceptual cue contributes to depth prediction.
By Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis
CamWorldQA introduces the first benchmark for assessing the perceptual quality of camera‑controlled world video generation, featuring 720 videos generated by six methods from 20 source videos across six camera trajectories, each scored by human raters. The paper also presents CWQA, a no‑reference quality assessment network that combines spatial, temporal motion, and optical flow features to predict quality scores. Experiments show CWQA outperforms existing VQA methods on the CamWorldQA dataset.
arXiv:2607. 00746v1 Announce Type: cross Abstract: The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception.
By Xiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang, Linna Song, Qingliang Luo, Chufan Guo, Kuifeng Su
The study examines how different two‑dimensional projections of spherical 360‑degree video affect end‑to‑end neural compression. Seven JVET‑360Lib formats were evaluated using a scale‑space flow model on standard test sequences, with performance measured by PSNR, spherical PSNR, weighted spherical PSNR, and BD‑rate. Results show that equirectangular and padded equirectangular projections yield the best compression efficiency with the neural model, while cubemap‑based formats excel with conventional HM‑16.16 codecs, highlighting that projection choice is codec‑dependent.
By Niloofar Maani
arXiv:2607. 12364v1 Announce Type: cross Abstract: EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning.
By Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik
O3N is a novel framework that performs open‑vocabulary occupancy prediction from a single omnidirectional RGB image. It introduces a polar‑spiral voxel embedding (PsM) for continuous 360° spatial representation, an Occupancy Cost Aggregation (OCA) module that unifies geometric and semantic supervision, and a Natural Modality Alignment (NMA) pathway that aligns visual, voxel, and text features. Experiments show state‑of‑the‑art results on QuadOcc and Human360Occ benchmarks, with strong cross‑scene generalization and semantic scalability.
By Mengfei Duan, Hao Shi, Fei Teng, Guoqiang Zhao, Yuheng Zhang, Zhiyong Li, Kailun Yang