The paper introduces OR‑Clarify, a benchmark that tests whether large language models can identify missing elements in natural‑language optimization requests before formulating a mathematical model. Each task provides a partial problem description and hides structured slots; agents interact with a simulated user to recover these slots, with metrics for accuracy, stopping decisions, and interaction cost. The authors also propose InterOPT, a two‑stage framework that detects unresolved gaps and decides whether to ask further questions or stop, achieving superior slot recovery in choice‑based experiments and competitive performance in open‑ended settings.
By Sihan Ge, Yichen Lin, Chenyu Zhou, Jianghao Lin, Tao Yao, Dongdong Ge
arXiv:2607. 20064v1 Announce Type: new Abstract: Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents.
By Alexis Fox, Junlin Wang, Paul Rosu, Bhuwan Dhingra
arXiv:2607. 02686v1 Announce Type: new Abstract: Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance from small language models (SLMs) that carry broad reasoning priors.
By Juarez Monteiro, Nathan Gavenski, Guilherme Lima, Francisco Galuppo, Odinaldo Rodrigues, Adriano Veloso
arXiv:2606. 16360v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting improves reasoning in large language models (LLMs) by externalizing intermediate computation as discrete text tokens, but this textual interface also introduces redundancy and inference overhead.
By Hanyu Lin, Min Cai, Jiawei Wen, Haodi Zhang
arXiv:2606. 03965v1 Announce Type: cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control.
By Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
arXiv:2607. 14105v1 Announce Type: cross Abstract: For Large Language Models to reliably answer user queries, users must clearly specify requirements, context, and constraints.
By Cedric Richter, Salah Ghamizi, Mike Papadakis
arXiv:2610.01892v1 Announce Type: cross
Abstract: Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning...
By Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
arXiv:2606. 16432v1 Announce Type: cross Abstract: User instructions are often underspecified because humans rely on implicit assumptions about the surrounding environment.
By Lai Jiang, Cheng Qian, Zhenhailong Wang, Pan Lu, Heng Ji, Hao Peng
arXiv:2608. 07968v1 Announce Type: cross Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time.
By Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou, Soheil Feizi
arXiv:2603. 02112v2 Announce Type: replace Abstract: Modern language models reason within bounded context, an inherent constraint that poses a fundamental barrier to long-horizon reasoning.
By Chenxiao Yang, Nathan Srebro, Zhiyuan Li
The paper introduces CLUE, a framework that lets robots actively resolve contextual uncertainty for underspecified natural language tasks. CLUE employs an LLM-derived policy to generate task-relevant hypotheses and plans, then uses an online language-embedded map to ground these into actions, refining its plan through closed-loop interaction. Experiments on a Boston Dynamics Spot across diverse indoor and outdoor settings show CLUE achieving near-oracle performance and outperforming LLM planners without closed-loop feedback by a significant margin.
By Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas, Vijay Kumar
JevSpawn is a new compositional policy that links natural language task specifications to finite probabilistic exploration, enabling LLM agents to generate actions more efficiently. It uses parallel action spawning, feedback‑driven branch selection, representation revision, and recovery from retained alternatives to adapt actions during interaction. Evaluations on eight benchmark tasks show that JevSpawn outperforms seven agent baselines and a TypeSafe Jev variant, improving task performance and speeding navigation.
By Haoyang Su, Weiran Huang