HEIR: Learning Human-Entity Interactions with Functional Roles
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.09736v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a partic...
arXiv:2606. 29613v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have recently been extended with role-based mechanisms for interpretability.
arXiv:2504.10079v5 Announce Type: replace Abstract: Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level repre...
The paper introduces the Identity-Aware Human-Object Interaction Motion Captioning task, which requires captions to include both the subject’s identity and the interaction motion, e.g., "Sub_ID lifts the chair" instead of a generic description. It proposes ID‑HOINet, a model that learns from multi‑view videos using a Multi‑View Identity‑Motion Learning Module and a Two‑Stage Caption Rewriting Strategy to generate identity‑aware captions. Experiments show that ID‑HOINet achieves state‑of‑the‑art performance on the BEHAVE and InterCap datasets.
arXiv:2608. 10765v1 Announce Type: new Abstract: Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy.
CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.