arXiv Computer Vision

NBAvatar: Neural Billboards Avatars with Realistic Hand-Face Interaction

NBAvatar is a method for realistic rendering of head avatars that handles non‑rigid deformations caused by hand‑face interaction. It introduces a hybrid implicit‑explicit representation, combining explicit oriented planar primitives with implicit neural rendering, and uses a geometry‑aware training scheme to jointly optimize these representations. The approach achieves up to 53% LPIPS reduction compared to Gaussian‑based avatar methods, improves PSNR and SSIM, and surpasses the state‑of‑the‑art InteractAvatar in structural similarity for novel‑view and novel‑pose rendering.

arXiv Computer Vision
Sep 18

AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars

AGORA is a new framework that extends 3D Gaussian Splatting with a generative adversarial network to produce high‑fidelity, animatable 3D head avatars. It introduces a lightweight FLAME‑conditioned deformation branch that predicts per‑Gaussian residuals for identity‑preserving, fine‑grained expression control, and a dual‑discriminator training scheme that enforces expression fidelity. The system achieves real‑time inference at 250 FPS on a single GPU and, for the first time, CPU‑only animatable 3DGS avatar synthesis at ~9 FPS.

By Ramazan Fazylov, Sergey Zagoruyko, Aleksandr Parkin, Stamatis Lefkimmiatis, Ivan Laptev
arXiv Computer Vision
Sep 25

PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark

PHOSA introduces MVSign, the first multi‑view Chinese sign language dataset co‑designed with Deaf experts, featuring diverse gestures and rich annotations. The authors develop a hybrid fitting pipeline for accurate SMPL‑X annotation and propose a decoupled sign avatar representation that isolates body, head, and hand components, coupled with a motion‑aware sampling strategy to handle motion blur and balance gesture diversity. Experiments show high‑fidelity visual results on MVSign, especially in detailed hand and facial regions, and good generalization to in‑the‑wild monocular sign language videos.

By Haodong Wang, Hezhen Hu, Wengang Zhou, Houqiang Li
arXiv Computer Vision
4d ago

AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling

AESplat is a new pose‑free feed‑forward 3D Gaussian Splatting framework that improves rendering quality by decoupling view‑independent and view‑dependent appearance modeling. It directly extracts the base view‑independent appearance from input images and predicts higher‑order spherical harmonic coefficients with a shallow MLP that incorporates 3D‑aware inductive biases. Experiments on several datasets show AESplat outperforms state‑of‑the‑art methods, achieving up to 0.8 dB higher PSNR than NAS3R and 1.1 dB over DepthSplat on RealEstate10K.

By Shiwei Ren, Zhiang Liu, Yongchun Fang, Hongwei Chen