arXiv AI

Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models

arXiv:2607. 24562v1 Announce Type: new Abstract: Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style.

arXiv Machine Learning
Aug 4

Conformalized Large Language Models under Configuration Shift

arXiv:2608. 01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability.

By Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang, Steffen Staab, Puneet Dokania, Philip Torr, Jie Tang, Evgeny Kharlamov
arXiv Machine Learning
1d ago

SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions

SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.

By Liang You, Hengyu Shi, Dongwen Ou
arXiv AI
2d ago

Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?

The paper introduces a benchmark for evaluating whether off‑the‑shelf small language models (SLMs) can reliably perform microtasks that support a large language model (LLM) planner, such as auto‑approving shell commands, writing memory, selecting tools, and ranking past turns. Using fixed prompts and confidence‑interval‑aware eligibility thresholds, the authors test several Qwen3 models (0.6/1.7/4/8 B) in FP16 with no tuning and find that none of the 16 configurations meet the eligibility criteria. Quantization to 4‑bit precision further degrades performance, with the eligibility gap tracking model size rather than precision, and the issue persists across different models (e.g., Llama‑3.x) and prompt variations.

By Jundong Hu, Shekar Ramachandran
arXiv Machine Learning
Aug 7

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.

By Ayushi Agarwal
arXiv AI
Sep 4

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

ESPO (Error-Structured Prompt Optimization) addresses prompt bloat in evolutionary prompt optimizers by splitting the optimization process into Diagnose, Propose, and Select phases. It clusters training errors into structural patterns, generates diverse candidate prompts through four complementary strategies, and applies bootstrap stability selection. Across seven NLP benchmarks, ESPO improves average accuracy by +3.76 pp over GEPA, produces prompts 47 % shorter, and achieves higher accuracy on four additional student models, with the largest gain on Qwen3 GSM8K.

By Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar
Hugging Face Trending Papers
Sep 3

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

ESPO (Error-Structured Prompt Optimization) addresses prompt bloat in evolutionary prompt optimizers by separating optimization into Diagnose, Propose, and Select phases. It clusters training errors, generates diverse candidates, and applies bootstrap stability selection, achieving a 3.76‑point accuracy gain over GEPA on seven NLP benchmarks while producing 47% shorter prompts. Cross‑model tests on four additional student models confirm ESPO’s superior average accuracy, notably improving Qwen3 GSM8K from 15.00% to 91.40%.