arXiv:2609.34187v2 Announce Type: replace-cross
Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
By Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack
arXiv:2608. 16068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks.
By Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao
The paper introduces ramework, a blackbox prompt‑minimization framework that identifies the minimal subset of few‑shot prompts necessary for large language models (LLMs). In a case study, the framework reduces few‑shot exemplars by an average of 65.3% in character count while maintaining full propositional output fidelity, revealing that models tend to keep logical identifiers and constraint declarations while discarding natural language prose. The analysis further distinguishes between universal encoder and decoder models, offering insights into prompt compression and structural analysis.
By Ali Alfageeh, Rahul Gopinath, Amin Alipour
arXiv:2607. 20500v1 Announce Type: new Abstract: Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer.
By Sizhe Tang, Guangyu Jiang, Yu Li, Rongqian Chen, Ioannis G. Kevrekidis, Tian Lan
arXiv:2608. 06933v1 Announce Type: cross Abstract: Today, we improve models by training and evaluating them on problems at the frontier of their abilities.
By Sarah Pratt, Jae Sung Park, Scott Geng, Ali Farhadi
arXiv:2608. 12426v1 Announce Type: new Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas.
By Mariya I. Vasileva
arXiv:2607. 20520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failures.
By Sagnik Nath, Edith Aurora Graf, Liang Zhang, Diego Zapata-Rivera
arXiv:2604.27251v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicite...
By Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Yuxiang Zhou, Maria Liakata, Nikolaos Aletras
The paper introduces OR‑Clarify, a benchmark that tests whether large language models can identify missing elements in natural‑language optimization requests before formulating a mathematical model. Each task provides a partial problem description and hides structured slots; agents interact with a simulated user to recover these slots, with metrics for accuracy, stopping decisions, and interaction cost. The authors also propose InterOPT, a two‑stage framework that detects unresolved gaps and decides whether to ask further questions or stop, achieving superior slot recovery in choice‑based experiments and competitive performance in open‑ended settings.
By Sihan Ge, Yichen Lin, Chenyu Zhou, Jianghao Lin, Tao Yao, Dongdong Ge
arXiv:2608.29610v1 Announce Type: new
Abstract: The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. W...
By Chenghao Yang
arXiv:2602. 15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibility--correctness gap reaching 90 percentage points on compositional problems.
By Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
arXiv:2606. 28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs.
By Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath