K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
arXiv:2607. 02680v1 Announce Type: cross Abstract: MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text.
The paper introduces ProAction, a multimodal dataset of 10,000 samples comprising visual, audio, and text inputs across 12 daily-life scenarios, designed to support the Proactive Robot Action Reasoning (ProRobo) problem. It presents a two-stage human-in-the-loop annotation pipeline that incorporates appraisal and Theory-of-Mind considerations to generate cognitively grounded high-level action labels. The authors benchmark multimodal large language models and propose MMC2Act, showing that training on ProAction significantly improves proactive action reasoning compared to general-purpose models.
arXiv:2607. 02680v1 Announce Type: cross Abstract: MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text.
The paper introduces a multimodal in‑context learning framework that uses contrastive demonstration modeling to align large language models’ responses with the required reasoning paths. By contrasting suboptimal and better responses and incorporating a response‑conditioned retrieval mechanism, the method explicitly guides models beyond surface imitation. Experiments on various multimodal tasks, especially visual question answering, show consistent performance gains.
arXiv:2608.21022v1 Announce Type: new Abstract: Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little consciou...
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task ca...
arXiv:2604. 00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention.
Zero-WAM introduces a causal video-action model that enables robots to perform unseen manipulation tasks by following in-context human video guidance. The authors create HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks, and propose an in-context future chunk prediction objective to prevent shortcut learning. In simulation, Zero-WAM attains a 47.0% success rate on seven unseen tasks, outperforming the best video-action baseline by 29.5 percentage points, and demonstrates real‑world generalization to complex, long‑horizon, and fine‑grained tasks.
arXiv:2512. 03438v3 Announce Type: replace Abstract: Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-based rewards computed based on the final answers.
arXiv:2607. 02542v1 Announce Type: new Abstract: General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons.
arXiv:2606. 15160v1 Announce Type: cross Abstract: Reasoning capabilities of multimodal large language models (MLLMs) have improved considerably in recent years.
The paper introduces Cognitive Chain-of-Thought (CoCoT), a structured reasoning framework for vision‑language models that divides multimodal social reasoning into three cognitively inspired stages: Perception, Situation, and Norm. CoCoT improves performance across diverse tasks—multimodal intent disambiguation, theory of mind, social commonsense reasoning, and safety instruction following—by 5.9% to 4.6% on average. Fine‑tuning on CoCoT‑structured traces further boosts accuracy by 5–6% without explicit prompting, indicating that models internalize the structured reasoning pattern and that the approach enhances interpretability and social alignment in multimodal systems.
ReactHuman is a physics‑grounded benchmark that tests whether multimodal large language models (MLLMs) can make immediate, safety‑critical decisions in simulated humanoid scenarios involving sudden household hazards. The benchmark includes 17 event families, over 1,000 reproducible scenes generated from 240 Hz rigid‑body simulation, and a five‑metric suite evaluating reactions on reasonableness, safety, and physical grounding. Evaluation of seven MLLMs reveals that reactive safety remains unsolved, with models frequently mishandling hazards, relying on appearance over motion, and missing key interception points.
arXiv:2609.36416v1 Announce Type: cross Abstract: Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions....