arXiv:2506. 07223v2 Announce Type: replace Abstract: Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safety-critical environments.
By Yangqing Zheng, Shunqi Mao, Dingxin Zhang, Weidong Cai
The paper introduces Harness Primitives—reusable agent harness mechanisms mined from failed task trajectories—and a framework called STITCH that selects and composes these primitives into task‑specific harnesses at test time. This approach avoids generating or debugging harness code for each task, achieving up to 12‑point gains in task success over fixed harness baselines and outperforming human‑designed harnesses like Codex CLI. STITCH also demonstrates minimal test‑time overhead (2.7%) and scales efficiently with the size of the primitive library.
By Peng Kuang, Haibo Jin, Dehao Wu, Feiyang Deng, Xiaopeng Yuan, Jerry Wang, Haohan Wang
arXiv:2606. 07999v1 Announce Type: new Abstract: Effective skill grounding is essential for deploying reusable skills in embodied agents, as even minor embodiment or environmental differences can render an entire skill incompatible.
By Sera Choi, Wonje Choi, Saehun Chun, Daehee Lee, Jooyoung Kim, Chaeun Lee, Honguk Woo
arXiv:2606. 14672v1 Announce Type: new Abstract: Large language models increasingly serve as execution engines for agentic systems, yet they still consume context through a sequential text interface.
By Shikun Liu, Mufei Li, Dongqi Fu, Haoyu Wang, Yinglong Xia, Hong Li, Hong Yan, Pan Li
arXiv:2606. 10087v1 Announce Type: cross Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats.
By Ankit Gupta, Aditya Prasad, Rameswar Panda
CEDAR is a counterexample-guided framework that translates natural-language instructions for embodied agents into regular languages over environment event traces, represented as deterministic finite automata. By using a language model for semantic judgments and execution traces for correction, CEDAR turns constraints into executable finite-state objects, enabling the intersection of learned skills with additional specifications. In Minecraft experiments, CEDAR preserves temporal and spatial constraints better than a program-generating baseline and reduces cumulative LLM queries by reusing learned skills.
By Lekai Chen, Alvaro Velasquez, Ashutosh Trivedi