arXiv Computation and Language

Beyond Correctness: Resolving Underspecification in Agentic Text-to-SQL

The paper introduces PlanPool, a method for agentic Text-to-SQL systems to manage clarification questions by maintaining a mutable question pool. By requiring agents to explicitly ask or drop each planned question, PlanPool improves ambiguity coverage and reduces silent failures compared to unconstrained or prompt-based approaches. Experiments on benchmarks derived from BIRD-Interact and Spider show that PlanPool achieves competitive execution accuracy while better handling underspecification.

arXiv Computation and Language
Sep 22

XYEval: Agents say yes to bad advice

arXiv:2609.23939v1 Announce Type: new Abstract: Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where...

By Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord, Been Kim
arXiv AI
2d ago

CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

CONTRA is a training‑free method that discovers and qualifies behavior‑changing questions for selective clarification in large language model (LLM) code generation. It first generates candidate questions, filters out those unrelated to required behavior or already resolved, then creates programs conditioned on two plausible answers to check for stable behavioral differences on shared inputs. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 across four coding agents, outperforming baselines by 13.88 percentage points, and it is also implemented as a Claude Code plugin for practical use.

By Zheng Fang, Yongmin Li, Yichang Zhang, Dongming Jin, Haoyu Wang, Shuai Wang, Zhi Jin, Ge Li
arXiv Computation and Language
Sep 3

A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models

The paper presents a tri‑agent framework for evaluating large language models’ question‑clarification abilities. It involves a Question Clarifying Agent that identifies ambiguities and asks follow‑up questions, a Respondent Agent that simulates human replies, and an Evaluator Agent that judges the dialogue using metrics such as ambiguity handling, question quality, dialogue efficiency, language appropriateness, and intent alignment. The authors illustrate the approach with synthetic supply‑chain data and discuss validating the evaluator against human judgments.

By Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra
arXiv AI
Aug 28

Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI

The paper argues that large language models need adaptive reasoning rather than fixed reasoning budgets. It shows that over‑reasoning leads to high computational cost without accuracy gains, while under‑reasoning results in incorrect or incomplete solutions. The authors evaluate these failure modes on MATH‑500 and the GAIA benchmark, highlighting the need for dynamic reasoning allocation in agentic AI systems.

By Md Jueal Mia, M. Hadi Amini
arXiv AI
Sep 7

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

The paper introduces OR‑Clarify, a benchmark that tests whether large language models can identify missing elements in natural‑language optimization requests before formulating a mathematical model. Each task provides a partial problem description and hides structured slots; agents interact with a simulated user to recover these slots, with metrics for accuracy, stopping decisions, and interaction cost. The authors also propose InterOPT, a two‑stage framework that detects unresolved gaps and decides whether to ask further questions or stop, achieving superior slot recovery in choice‑based experiments and competitive performance in open‑ended settings.

By Sihan Ge, Yichen Lin, Chenyu Zhou, Jianghao Lin, Tao Yao, Dongdong Ge