The Planning Limits of Latent World Models
arXiv:2609.39235v1 Announce Type: cross Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.
arXiv:2609.39235v1 Announce Type: cross Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
arXiv:2609.39964v1 Announce Type: new Abstract: Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. D...
arXiv:2609.38383v1 Announce Type: cross Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without po...
arXiv:2609.38536v1 Announce Type: cross Abstract: Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding,...
arXiv:2609.40137v1 Announce Type: cross Abstract: We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans...
GestAdapt is a framework that generates co‑speech gestures conditioned on a specified wrist workspace, allowing humanoid robots to adapt their motions to environmental constraints such as walls. The system learns from six co‑speech corpora using a shared motion representation and can be retargeted to different robot embodiments. Experiments show that GestAdapt’s motions stay close to real‑motion distributions, achieve higher quality scores than a no‑workspace baseline, and outperform other methods in real‑robot evaluations on the Reachy2 humanoid.
arXiv:2609.38653v1 Announce Type: cross Abstract: Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human mo...
arXiv:2609.38982v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital w...
arXiv:2511.10661v2 Announce Type: replace Abstract: It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often re...
The paper introduces FIGS, a dual‑axis evaluation framework for multi‑turn sycophancy that avoids penalizing empathy. It uses a 10‑turn conversational simulator with 500 diverse scenarios to test whether models stay truthful while keeping praise proportional, and whether they show calibrated validation of user feelings. The study finds that current models either drift toward sycophancy or become overly detached, highlighting an unresolved trade‑off in sustained dialogue.
DyRAD introduces a novel radar novel‑view synthesis framework that models dynamic driving scenes by separating static background reflectors from motion‑tracked dynamic point reflectors, enabling the rendering of full range‑azimuth‑Doppler (RAD) tensors. The method derives reflector velocities from object tracks, projects them onto the line of sight, and uses a fixed analytic point‑spread function to avoid embedding sensor‑induced spread into the scene representation. This design allows accurate scene reconstruction and zero‑shot transfer to different radar configurations, achieving a 90.7% recovery of radar detections on the RADIal dataset compared to 26.9% for the best baseline.
The paper investigates how generative behavioral-cloning policies handle multimodal expert behavior, identifying bottlenecks in both latent-variable and action-space policy designs. For latent-variable policies, preserving demonstrated modes depends on action-conditioned latent representations, and excessive posterior-prior regularization can suppress this information. In action-space generative policies, multimodality is limited by the smoothness of the base-to-action transport, requiring either sharp transitions or off-support bridge regions to capture many well-separated modes. Experiments on synthetic navigation and a physical-robot manipulation task confirm these findings, while standard robotic simulation benchmarks show limited conditional multimodality, making deterministic regression competitive.
The paper introduces EvolvingNav, a system that builds a time‑indexed belief about moving targets in dynamic environments by combining timestamped 3D object histories with a persistence‑relocation model. It uses an event‑driven filter to update beliefs over time, incorporates RGB‑D evidence, and applies a zero‑shot vision‑language controller for action selection. The authors also present EvoWorld‑Bench, a large benchmark of human‑activity‑based scenes, and demonstrate that EvolvingNav outperforms baselines in both simulation and real‑robot experiments, especially when temporal patterns are learnable.
DiFF is a generative framework that uses Doppler velocity cues from 4D millimeter-wave radar to improve human motion flow estimation. It combines Doppler-informed motion priors with a Kolmogorov‑Arnold Network (KAN) based conditional flow matching model, featuring a KAN‑attention mechanism for expressive feature extraction. Experiments demonstrate that DiFF achieves state‑of‑the‑art performance, reducing 3D endpoint error to the millimeter scale on the mmBody benchmark.
arXiv:2609.39378v1 Announce Type: new Abstract: Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolvin...
ECHO-G is a framework for generating full‑body co‑speech motion for humanoid robots, jointly conditioned on speech audio and timed transcripts. Its Speech‑Grounded Diffusion Transformer (SGDiT) fuses frame‑aligned acoustic features with token‑level linguistic context, preserving distinct granularities while modeling one‑to‑many utterance‑motion relationships directly in robot space. The authors introduce a BEAT2‑derived audio‑text‑robot dataset, a benchmark for co‑speech characteristics, robot‑motion quality, and runtime efficiency, and demonstrate that direct robot‑space generation outperforms human‑motion generation and retargeting pipelines, with joint audio‑text conditioning yielding superior results in both quantitative evaluation and a video‑rating study. "whyItMatters":"The study provides a new dataset, benchmark, and a demonstrably effective method for generating realistic co‑speech motion directly in robot space, advancing practical humanoid robot interaction."
Ego4WAM investigates how various properties of egocentric human data—such as human‑robot alignment, data duration, task diversity, and supervision type—affect robot learning. The study, conducted under a unified world‑action model framework, shows that aligned demonstrations improve out‑of‑distribution generalization and lower the amount of target‑task robot data needed. It also finds that video‑only supervision remains effective, and that data duration and task diversity influence downstream capabilities in distinct ways, as validated on real robots and RoboDojo.
SYNCR is a synthetic benchmark designed to evaluate multimodal large language models on cross‑video reasoning. It contains 4,000 question‑answer pairs across 4,827 unique videos, covering tasks in temporal alignment, spatial tracking, comparative reasoning, and holistic synthesis. The benchmark reveals a significant performance gap between current models and humans, with models excelling at temporal ordering but struggling with precise physical and spatial reasoning.
SkillFM is a generative framework that creates task‑conditioned textual skills for large language model agents without relying on manual skill banks or reinforcement learning. It encodes skills into a continuous latent space using a codec and trains a conditional flow model with improved MeanFlow, allowing single‑step latent sampling at inference. The sampled latent is decoded by an LLM into textual guidance, and the method outperforms other vector‑based skill approaches on ALFWorld, Search‑QA, and other tasks.
SynIL is a new framework for offline imitation learning that automatically assesses the quality of demonstration data without requiring labels. It uses motor synergy—a low‑dimensional coordinated movement pattern linked to proficiency—to generate dense, transition‑level reward signals through self‑supervised reward regression. Experiments on D4RL locomotion and Robomimic manipulation datasets show that synergy‑derived rewards align well with true rewards and that SynIL outperforms Behavior Cloning and rivals or surpasses offline reinforcement learning in sparse‑reward scenarios.