UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
arXiv:2608. 03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet.
arXiv:2608. 08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.
arXiv:2608. 03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet.
arXiv:2607. 03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, robot states, human demonstrations, and task contexts into structured representations, decisions, trajectories, control references, and system services.
arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.
arXiv:2606. 11909v1 Announce Type: new Abstract: Benchmarks are essential for evaluating embodied spatial intelligence, yet their construction is labor-intensive, hard to reuse, and difficult to maintain.
arXiv:2609. 30971v1 Announce Type: new Abstract: Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments.
arXiv:2510.26915v2 Announce Type: replace-cross Abstract: While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large...
arXiv:2507. 07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate these skills simultaneously.
AeroWeaver is a new embodied‑agent harness that integrates large language model (LLM) decision making with the executable skills of individual UAVs, enabling distributed, adaptive swarm execution. It connects semantic mission decisions to governed skills, organizes role‑conditioned local agents for coordination, and refines skill selection online using role‑indexed state‑action‑reward experience. Experiments demonstrate that AeroWeaver maintains valid skill execution without a central joint‑action generator and supports reward‑guided, training‑free adaptive learning from accumulated execution experience.
arXiv:2609.24526v1 Announce Type: new Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and...
MineExplorer is a benchmark designed to assess the open‑world exploration abilities of multimodal large language models (MLLMs) in Minecraft. It filters out tasks that rely heavily on Minecraft‑specific knowledge, organizes tasks into ReAct‑style capabilities, and composes atomic tasks into implicit multi‑hop challenges. A multi‑agent synthesis workflow creates reliable task graphs, sandbox scenes, and rule‑based milestone evaluators, and human evaluation confirms its superiority over a single‑agent baseline. Experiments show that while advanced MLLMs can handle many single‑hop tasks, they struggle with longer trajectories that require coordinating hidden prerequisites, and larger models or different thinking modes do not consistently improve performance.
arXiv:2609.24526v2 Announce Type: replace Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints...
AURORA is a natural‑language‑driven framework that treats air‑ground scenario generation as a compilation process with verification. It introduces the Air‑Ground Scenario Graph (AGSG), a typed intermediate representation linking agents, missions, events, communication, and success conditions, enabling joint grounding, temporal planning, pre‑execution checks, runtime verification, failure localization, and bounded repair. The authors also present AURORA‑Bench to evaluate not only execution but faithful realization of requested interactions, showing that structured execution and runtime verification improve reliability and that explicit intermediate representations facilitate verifiable and repairable co‑simulation.