EgoForge: Goal-Directed Egocentric World Simulator
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2610.01092v1 Announce Type: new Abstract: Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must...
arXiv:2604. 01001v2 Announce Type: replace-cross Abstract: We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation.
MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.
The paper introduces a human‑centric video world model that can be controlled by tracked head pose and joint‑level hand poses, enabling realistic hand‑object interactions. It evaluates existing diffusion transformer conditioning methods and proposes a new 3‑D control mechanism, training a bidirectional diffusion model and distilling it into a causal, interactive system for egocentric virtual environments. Human‑subject experiments show that the system improves task performance and gives users a higher perceived sense of control compared to baseline methods.
arXiv:2607. 17790v1 Announce Type: cross Abstract: Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment.
arXiv:2608.20308v2 Announce Type: replace Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe ob...