Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
arXiv:2608. 08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.
AURORA is a natural‑language‑driven framework that treats air‑ground scenario generation as a compilation process with verification. It introduces the Air‑Ground Scenario Graph (AGSG), a typed intermediate representation linking agents, missions, events, communication, and success conditions, enabling joint grounding, temporal planning, pre‑execution checks, runtime verification, failure localization, and bounded repair. The authors also present AURORA‑Bench to evaluate not only execution but faithful realization of requested interactions, showing that structured execution and runtime verification improve reliability and that explicit intermediate representations facilitate verifiable and repairable co‑simulation.
arXiv:2608. 08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.
MOOSEnger is a simulation‑aware AI agent framework designed for the MOOSE ecosystem, integrating an interchangeable reasoning model with domain knowledge, revised simulation artifacts, MOOSE‑specific validation, and executable solver feedback. Its generate‑check‑repair‑run workflow uses MOOSE knowledge retrieval, HIT‑aware parsing, syntax metadata, diagnostics, and revision‑controlled authoring to bind evidence to each input revision and guide bounded repair before acceptance. Across 200 prompts, MOOSEnger raises executable success from 5% to 89.5% with GPT‑5.2 and from 0% to 76.5% with Gemma 4 31B, and a ten‑case benchmark shows all generated inputs meet semantic alignment, with eight also meeting numerical‑accuracy criteria.
arXiv:2608. 16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of...
arXiv:2610.01093v1 Announce Type: cross Abstract: Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process in which engineers translate high-l...
arXiv:2604. 02478v2 Announce Type: replace Abstract: Deep learning models excel at detecting anomaly patterns in normal data.
arXiv:2608. 03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet.
arXiv:2510. 12985v3 Announce Type: replace Abstract: We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents.
arXiv:2603. 26005v2 Announce Type: replace Abstract: Grid-interactive building control has emerged as a promising approach for improving demand-side flexibility in modern power systems.
arXiv:2505. 18334v2 Announce Type: replace-cross Abstract: Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other.
arXiv:2508. 02721v2 Announce Type: replace-cross Abstract: While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and predictable execution are strict requirements.
arXiv:2607. 16266v1 Announce Type: cross Abstract: Existing approaches for reasoning about action and change provide expressive semantics for modeling dynamic systems, in most cases built on top of logic programming systems.