Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
arXiv:2606. 19980v1 Announce Type: new Abstract: Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence.
FluxVLA Engine is an open, configuration‑driven platform that unifies the fragmented components of embodied policy development—datasets, visual‑language and world models, action heads, learning methods, distributed training, simulation evaluation, inference, and robot interfaces—into a reproducible data‑to‑deployment workflow. It adds features such as compositional dual‑arm simulation, scalable automatic data generation, human‑in‑the‑loop rollout and correction, Real‑Time Chunking for fast inference, and lightweight remote GPU serving, thereby linking offline learning, simulation validation, online correction, and real‑robot execution under shared, auditable contracts. The engine aims to eliminate engineering bottlenecks that currently separate promising embodied‑learning algorithms from reliable, reproducible deployment.
arXiv:2601. 20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task-specific demonstrations and fine-tuning, and often generalize poorly under domain shift.
arXiv:2607. 11689v1 Announce Type: cross Abstract: Artificial general intelligence ultimately requires agents that can reason and act in the physical world.
arXiv:2607. 00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures.
arXiv:2608.29596v1 Announce Type: new Abstract: Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on c...
The paper introduces SHAPER, a self‑evolving framework that enables embodied agents to adapt to new environments without retraining the underlying foundation model. SHAPER keeps model parameters frozen and instead evolves reusable skills and a context‑code harness through rollouts in the target environment, allowing the same model to act as both planner and optimizer. Experiments on VLABench and ESI‑Bench demonstrate that this skill‑and‑harness optimization outperforms pure execution, supervised fine‑tuning, and test‑time scaling baselines, showing that self‑evolving adaptation is viable when model training is costly or impractical.
arXiv:2607. 22832v1 Announce Type: new Abstract: Long-horizon embodied tasks require policies that execute many dependent actions before task success can be observed.
arXiv:2607. 03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, robot states, human demonstrations, and task contexts into structured representations, decisions, trajectories, control references, and system services.
arXiv:2609. 30971v1 Announce Type: new Abstract: Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments.
arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.