arXiv Computer Vision By Weikun Peng, Denys Iliash, Manolis Savva

EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates

Read the original on arXiv Computer Vision →

EgoFun3D introduces a coordinated task, dataset, and benchmark for creating simulation-ready interactive 3D objects from egocentric videos. The approach captures general cross-part functional mappings via function templates, enabling precise evaluation and direct compilation into executable code. A 4‑stage pipeline—2D part segmentation, reconstruction, articulation estimation, and function template inference—is proposed, and a dataset of 517 videos with detailed annotations is released for comprehensive benchmarking.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jul 2

EgoSim: Egocentric World Simulator for Embodied Interaction Generation

arXiv:2604. 01001v2 Announce Type: replace-cross Abstract: We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation.

By Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu, Jiafei Cao, Xihui Liu, Ran Yi, Lizhuang Ma, Jiangmiao Pang, Xudong Xu
arXiv Computer Vision
6d ago

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Ego2Act is a new benchmark that tests video generation models on goal‑directed, egocentric manipulation tasks. It contains 2,640 videos from 110 real‑world tasks, each requiring multiple steps of object manipulation to achieve a high‑level goal. The benchmark includes Ego2ActJudge, a reference‑free evaluation pipeline that better aligns with human judgments of task completion and physics plausibility.

By Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, Alham Fikri Aji
arXiv AI
Jun 29

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.

By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha
arXiv AI
Sep 18

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

FAMOS is a feed‑forward model that predicts movable‑part segmentation and joint parameters from a sparse, unordered set of partial point clouds. It jointly reasons over multiple observations using a Multi‑state Articulation Transformer that alternates state‑wise and global attention, and introduces an observed articulation span objective to supervise motion ranges across inputs. A procedural data generator supplies self‑annotated assets for training, and experiments on PartNet‑Mobility, ACD, and ArtiCraft‑10K show consistent improvements over existing feed‑forward and optimization‑based baselines.

By Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni
arXiv Computer Vision
6d ago

EgoForge: Goal-Directed Egocentric World Simulator

arXiv:2603.20169v2 Announce Type: replace Abstract: Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint cha...

By Yifan Shen, Jiateng Liu, Xinzhuo Li, Yuanzhe Liu, Bingxuan Li, Houze Yang, Wenqi Jia, Yijiang Li, Tianjiao Yu, James Matthew Rehg, Xu Cao, Ismini Lourentzou