arXiv:2606. 29054v1 Announce Type: new Abstract: Large language models (LLMs) deployed for structured generation (NER, JSON extraction, QA, and classification) lack formal reliability guarantees, and standard heuristic abstention policies miss user-specified risk targets by 7.
By Varun Kotte
arXiv:2608. 05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human.
By Jianru Shen
arXiv:2608. 01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability.
By Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang, Steffen Staab, Puneet Dokania, Philip Torr, Jie Tang, Evgeny Kharlamov
SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.
By Liang You, Hengyu Shi, Dongwen Ou
The paper introduces a benchmark for evaluating whether off‑the‑shelf small language models (SLMs) can reliably perform microtasks that support a large language model (LLM) planner, such as auto‑approving shell commands, writing memory, selecting tools, and ranking past turns. Using fixed prompts and confidence‑interval‑aware eligibility thresholds, the authors test several Qwen3 models (0.6/1.7/4/8 B) in FP16 with no tuning and find that none of the 16 configurations meet the eligibility criteria. Quantization to 4‑bit precision further degrades performance, with the eligibility gap tracking model size rather than precision, and the issue persists across different models (e.g., Llama‑3.x) and prompt variations.
By Jundong Hu, Shekar Ramachandran
arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.
By Ayushi Agarwal
arXiv:2608. 14509v1 Announce Type: new Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt.
By Zhelun Wu
arXiv:2609.24303v1 Announce Type: new
Abstract: Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident t...
By Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, Jos\'e Miguel Hern\'andez-Lobato, Junbin Gao
ESPO (Error-Structured Prompt Optimization) addresses prompt bloat in evolutionary prompt optimizers by splitting the optimization process into Diagnose, Propose, and Select phases. It clusters training errors into structural patterns, generates diverse candidate prompts through four complementary strategies, and applies bootstrap stability selection. Across seven NLP benchmarks, ESPO improves average accuracy by +3.76 pp over GEPA, produces prompts 47 % shorter, and achieves higher accuracy on four additional student models, with the largest gain on Qwen3 GSM8K.
By Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar
arXiv:2607. 23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B".
By Paul Simpson, John Kozak, Lisa Doake
ESPO (Error-Structured Prompt Optimization) addresses prompt bloat in evolutionary prompt optimizers by separating optimization into Diagnose, Propose, and Select phases. It clusters training errors, generates diverse candidates, and applies bootstrap stability selection, achieving a 3.76‑point accuracy gain over GEPA on seven NLP benchmarks while producing 47% shorter prompts. Cross‑model tests on four additional student models confirm ESPO’s superior average accuracy, notably improving Qwen3 GSM8K from 15.00% to 91.40%.
arXiv:2607.27836v2 Announce Type: replace
Abstract: Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recover...
By Xiangyu Yin, Jiaxu Liu, Zhen Chen, Chih-Hong Cheng