Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms.
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.
arXiv:2609.38347v1 Announce Type: new Abstract: Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visuall...
ORMA is a training‑free framework that reconstructs articulated 4D representations of animals from monocular videos by decoupling pose and shape. It uses predicted pose as a reference for optimization and generative 3D priors to refine shape, aligning the result with the SMAL+ parametric model. The method combines per‑frame pose estimates with globally consistent camera poses, and further refines the reconstruction using self‑supervised DINO correspondences and temporal consistency, achieving improved accuracy on the new PAW4D benchmark and diverse real‑world videos.
arXiv:2606. 07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces.
arXiv:2609.19142v1 Announce Type: new Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse...