Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
arXiv:2606. 28529v1 Announce Type: cross Abstract: Embodied foundation models have recently been widely used to improve robot generalization and task success rates.
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
The paper introduces rMuscle, a real‑time Vision‑Language‑Action inference framework that mimics human muscle memory to accelerate robotic decision making. By exploiting repeated task similarity, rMuscle uses a dual‑phase cache: a Context Cache reuses visual‑token outputs and an Action Cache reuses neuron activation patterns, reducing computation and weight accesses. Experiments on RTX 4090 and Jetson Thor show 1.29–1.42× speedups on LIBERO, RoboTwin, and physical manipulation tasks while preserving success rates on real robots.
Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies...
The paper introduces Interaction‑Aligned Pruning (IAprune), a training‑free method for visual token pruning in embodied manipulation tasks. IAprune jointly decides per‑frame budget and token selection, using semantic‑motion spatial agreement to choose between conservative and aggressive coverage, and applies geometric residual correction to focus on under‑represented boundaries. Experiments on four policies, three simulation benchmarks, and a real‑robot platform show that IAprune matches unpruned performance on LIBERO while achieving up to 1.54× speed‑up and 1.48× acceleration on a real robot.
arXiv:2608.22971v1 Announce Type: new Abstract: Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and intera...
arXiv:2512. 01031v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks.
The paper introduces ProgressCompass, a framework that enhances Embodied Progress Reward Models (PRMs) by providing the necessary contextual information for accurate progress estimation in long manipulation tasks. It presents ContextProgress-Bench, a benchmark with 24 tasks that tests PRMs under three context-dependent scenarios—State Recall, Sequence Tracking, and Recurrence Disambiguation—showing that even history-aware PRMs struggle without proper context. By integrating a context-aware loop that leverages general-purpose vision‑language models, ProgressCompass reduces PRM progress error by up to 82% and improves rank agreement by 76%.
arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.
arXiv:2511. 15669v3 Announce Type: replace-cross Abstract: Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead?
Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studie...
arXiv:2608.21247v1 Announce Type: new Abstract: Token compression has become a key technique for reducing the inference cost of large foundation models, with approaches such as token pruning and KV-c...
arXiv:2608. 06756v1 Announce Type: new Abstract: Vision-language models are increasingly serving as the reasoning core of embodied agents.