arXiv:2609.39378v1 Announce Type: new
Abstract: Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolvin...
By Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu
The paper introduces a compact framework that transforms continuous multimodal workplace video into a structured Procedural State Memory called a Work Environment Model (WEM). Using event segmentation theory, it detects segment boundaries based on changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, then abstracts each segment into an evidence‑linked event card. These event cards incrementally update the WEM, enabling efficient, auditable documentation and retrieval while respecting on‑premise privacy constraints, and the authors evaluate the system on segmentation quality, memory compression, retrieval fidelity, and long‑horizon QA.
By Vivek Chavan, J\"org Kr\"uger
arXiv:2608. 13684v1 Announce Type: new Abstract: This paper describes a neurosymbolic architecture for learning to assemble novel structures using evidence from embodied conversations and task demonstrations.
By Jonghyuk Park, Alex Lascarides, Subramanian Ramamoorthy
KnowDemo is a framework that generates diverse robot demonstrations from human videos by leveraging structured manipulation knowledge. It uses a vision‑language model to extract task requirements and permissible execution variations, then resolves these against target‑scene entities to guide candidate generation and screening before motion planning. The resulting demonstrations feature multimodal behavior, alternative contact strategies, and valid subtask orders, and have been shown to improve planning success and enable sim‑to‑real policy transfer across three tasks.
By Zhiyuan Gao, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Sch\"afer, Michael Beetz
arXiv:2606. 31980v1 Announce Type: cross Abstract: Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves?
By Meng Chen, Anya Ji, Tsung-Han Wu, Tobias Maringgele, David M. Chan, Alane Suhr, Amy Pavel
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the first systematic study of real-scale bimanual furniture assembly using Vision-Language-Action models (VLAs).