arXiv Computer Vision

DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients

Hugging Face Trending Papers
Aug 6

Floating Radiance Networks

Recent advances in neural scene representations enable photorealistic novel-view synthesis, yet most methods remain tightly coupled to a single rendering paradigm, limiting their versatility and integration with conventional graphics workflows. We introduce Floating Radiance Networks (FlaRe), a neural scene representation combining explicit ray-traceable geometry with continuous neural radiance functions.

Hugging Face Trending Papers
Jul 20

FF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera System

Projector-camera (ProCams) systems achieve active scene perception and controllable appearance manipulation via structured illumination, serving as a core infrastructure for spatial augmented reality, projection mapping, and surface reflectance acquisition. Existing inverse-rendering methods for ProCams deliver high-fidelity results but rely on time-consuming per-scene optimization, while mainstream feed-forward 3D reconstruction models produce baked appearance that cannot adapt to spatially varying projector illumination.

arXiv Computer Vision
Sep 2

Hydra: Marker-Free RGB-D Hand-Eye Calibration

Hydra introduces a marker‑free RGB‑D hand‑eye calibration method that leverages a novel ICP algorithm with a robust point‑to‑plane objective on a Lie algebra. Experiments on three serial manipulators and two RGB‑D cameras show that with only three random robot configurations the method achieves about 90% successful calibrations, 2–3× faster convergence to the global optimum, and 2 orders of magnitude faster convergence time (0.8 ± 0.4 s) compared to other marker‑free baselines. The approach delivers improved accuracy (5 mm in task space versus 7 mm for classical methods) while remaining marker‑free, and the authors provide an open‑source dataset, code, and ROS 2 integration.

By Martin Huber, Huanyu Tian, Christopher E. Mower, Lucas-Raphael M\"uller, S\'ebastien Ourselin, Christos Bergeles, Tom Vercauteren
arXiv Computer Vision
Sep 28

From Mono to Stereo: Accelerating Binocular Gaussian Splatting via Reprojection and Selective Patching

The paper introduces a 2D Gaussian Splatting pipeline that renders a dominant-eye RGB image and depth proxy, then reprojects and selectively patches the affiliated eye to reduce redundant work. By reusing alpha-blending weights and generating adaptive regions of interest, the method cuts sequential binocular rendering time by 15.5% to 28.8% and GPU memory by 6% to 11% on several datasets, with minimal quality loss. It demonstrates a practical efficiency‑quality trade‑off for static‑scene stereo rendering and suggests further evaluation on dynamic scenes and VR hardware.

By Hongfei Zhu, Ling Zhou
arXiv Computer Vision
Sep 18

BinoGen: Scaling egocentric binocular data for embodied visual perception and learning

BinoGen is an automated framework that generates large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. It models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups, producing synchronized videos with dense multimodal supervision such as depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, the authors created a dataset of over 20 million annotated images, demonstrating that incorporating this data improves real-world visual perception tasks like depth estimation, object detection, and video object tracking, and that embodiment-specific adaptation enhances performance while joint training enables a single model to perform competitively across different observer embodiments.

By Chunpeng Li, Ya-tang Li
arXiv Computer Vision
Sep 28

NBAvatar: Neural Billboards Avatars with Realistic Hand-Face Interaction

NBAvatar is a method for realistic rendering of head avatars that handles non‑rigid deformations caused by hand‑face interaction. It introduces a hybrid implicit‑explicit representation, combining explicit oriented planar primitives with implicit neural rendering, and uses a geometry‑aware training scheme to jointly optimize these representations. The approach achieves up to 53% LPIPS reduction compared to Gaussian‑based avatar methods, improves PSNR and SSIM, and surpasses the state‑of‑the‑art InteractAvatar in structural similarity for novel‑view and novel‑pose rendering.

By David Svitov, Mahtab Dahaghin, Pietro Morerio, Alessio Del Bue