STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
arXiv:2610.10438v1 Announce Type: new Abstract: Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment an...
arXiv:2604. 01001v2 Announce Type: replace-cross Abstract: We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation.
Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it. Progre...
MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.
The paper introduces a metric interaction framework for robotic manipulation that explicitly models object- and scene-level interactions in Cartesian space. It uses Interaction‑Centric Tokens (ICTs) to represent end‑effector trajectories relative to objects and a Metric Action Interaction Field (MAIF) to attend to scene point‑cloud features for geometry‑conditioned action corrections. Experiments show modest but consistent improvements across several benchmarks, including LIBERO, RoboTwin 2.0, and real‑world tasks.
arXiv:2607. 08436v1 Announce Type: cross Abstract: Egocentric human data offers scalable supervision for robot manipulation.
arXiv:2609.22332v1 Announce Type: cross Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act...
arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.
The paper introduces Hand2Bot, an RGB‑D video dataset designed for human‑to‑robot handover scenarios, capturing body posture and facial expressions amid real‑world noise. It also proposes PassGen, a generative pipeline using stable video diffusion and an Intention‑Aware Temporal Face Encoder to synthesize realistic handover sequences while maintaining hand‑object consistency. A morphology‑based depth editing strategy is employed to replicate realistic sensor noise, and experiments show that training on PassGen yields high intention identification accuracy, low false trigger rates, and robust zero‑shot transfer to a physical robot platform.
arXiv:2607.08098v2 Announce Type: replace Abstract: Event cameras are increasingly adopted in embodied perception for their microsecond temporal resolution, high dynamic range, and resilience to moti...
arXiv:2609.38443v1 Announce Type: cross Abstract: We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yi...