arXiv AI By Vivek Chavan, J\"org Kr\"uger

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

Read the original on arXiv AI →

This paper introduces a data‑centric method that extracts structured task knowledge from expert demonstrations in assembly and disassembly operations. By jointly encoding temporal and multimodal data from egocentric and exocentric video recordings and narration, the approach produces task representations that support procedural documentation and context‑aware worker guidance. Evaluation on a real‑world disassembly case study shows that video‑based representations capture procedural structure and execution context more effectively than static image‑based methods, underscoring the value of egocentric video understanding for repair, training, and circular manufacturing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
3d ago

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

arXiv:2609.39378v1 Announce Type: new Abstract: Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolvin...

By Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu
arXiv AI
2d ago

A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction

The paper introduces a compact framework that transforms continuous multimodal workplace video into a structured Procedural State Memory called a Work Environment Model (WEM). Using event segmentation theory, it detects segment boundaries based on changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, then abstracts each segment into an evidence‑linked event card. These event cards incrementally update the WEM, enabling efficient, auditable documentation and retrieval while respecting on‑premise privacy constraints, and the authors evaluate the system on segmentation quality, memory compression, retrieval fidelity, and long‑horizon QA.

By Vivek Chavan, J\"org Kr\"uger
arXiv AI
Sep 21

KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

KnowDemo is a framework that generates diverse robot demonstrations from human videos by leveraging structured manipulation knowledge. It uses a vision‑language model to extract task requirements and permissible execution variations, then resolves these against target‑scene entities to guide candidate generation and screening before motion planning. The resulting demonstrations feature multimodal behavior, alternative contact strategies, and valid subtask orders, and have been shown to improve planning success and enable sim‑to‑real policy transfer across three tasks.

By Zhiyuan Gao, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Sch\"afer, Michael Beetz