Hugging Face Trending Papers

BayesPrompt: human readable prompts that make sense

BayesPrompt proposes a Bayesian approach to prompt optimisation for large language models, aiming to generate prompts that are both efficient in perplexity and human readable. The authors argue that traditional optimisation methods produce unintelligible pseudoprompts due to the ill‑posed nature of the task. Their algorithm samples prompts from a posterior distribution, and experiments on real data show marked improvements over state‑of‑the‑art alternatives across several metrics.

arXiv AI
Jun 2

CA-BED: Conversation-Aware Bayesian Experimental Design

arXiv:2606. 01182v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at static reasoning tasks, yet their performance often degrades in interactive scenarios where information must be actively acquired through questioning.

By Daniel Arnould, Rashad Aziz, Zixuan Kang, Tanav Changal, Kevin Zhu, Sunishchal Dev, Gabriel Grand, Shreyas Sunil Kulkarni
arXiv Computation and Language
Sep 25

Likelihood Ranking doesn't Scale Like Prompting in LLMs

The paper compares two common ways of evaluating large language models (LLMs): prompting them to answer questions directly and scoring candidate answers using likelihood-based metrics. The authors introduce a new protocol that ranks declarative statements derived from question–answer pairs, and test it across 95 decoder-only models (0.1B–104B parameters) on 10 multiple-choice QA datasets. They find that while prompted answering accuracy improves sharply with model scale and instruction tuning, statement‑likelihood ranking accuracy stays relatively stable, indicating that the two evaluation methods probe different aspects of model behavior.

By Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci
arXiv Machine Learning
Sep 3

Prompting the Unknown: Understanding Response Uncertainty in Large Language Models

The paper introduces a prompt-response concept model that links the amount of task-relevant information in a prompt to the uncertainty of responses generated by large language models (LLMs). It identifies four sources of response uncertainty—prompt underspecification, model quality, task variability, and semantic redundancy—and demonstrates that uncertainty decreases as prompt informativeness or model quality increases, analogous to epistemic uncertainty in probabilistic models. Experiments on real-world datasets confirm the theoretical predictions and validate the model.

By Ze Yu Zhang, Arun Verma, Finale Doshi-Velez, Bryan Kian Hsiang Low
arXiv AI
4d ago

Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity

The paper investigates prompt minimization, aiming to reduce prompts to their smallest, most information-dense form without losing output fidelity. It argues that shorter prompts lower computational overhead and inference latency, especially when large contexts are unnecessarily included, and that longer prompts can harm LLM reasoning and accuracy. The authors propose three frameworks to identify minimal prompts and show that these often produce outputs comparable to longer versions, highlighting redundancy in the input space and opening new avenues for efficient prompt engineering.

By Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu
arXiv AI
Jul 7

Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs

arXiv:2607. 03426v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions required in sequential decision-making settings.

By Jakob Hartmann, James Harvey, Jhonathan Navott, Erik Y. Wang, Luckeciano C. Melo, Flaviu Cipcigan, Cheng Zhang, Alessandro Abate
arXiv Computation and Language
Sep 17

Reading Between the Lines: Can LLMs Discover the Question Behind the Text?

The paper introduces "question archaeology," an evaluation task that asks models to infer the single, authentic question that motivated a text. It presents a new dataset of commissioned texts paired with their original research questions and distractors, and evaluates both proprietary and open‑source LLMs. Results show newer models outperform older ones, with BERT-based models lagging, and current LLMs even surpassing human performance on this task.

By Claudiu Creanga, Liviu P. Dinu
arXiv AI
Sep 15

Bayesian Intelligence from the Outside

The paper introduces a Bayesian framework for assessing intelligence in agents such as language models. It shows that an agent’s reports are consistent with Bayesian intelligence if they are not fully contradictory across prompts, and it defines an intelligence order based on the informativeness of internal experiments. The work also demonstrates the challenges of aggregating coarse reports from intelligent agents, revealing that optimal aggregation can assign arbitrary weights to unexcluded states unless the agent reports a belief about the complete state of the world.

By Alex Smolin, Bryan Wilder
arXiv Computation and Language
Sep 2

PromptNCE: Conditional Probabilities and PMI Using Only LLMs and Contrastive Estimation Prompts

The paper introduces PromptNCE, a zero‑shot method that uses large language models to estimate pointwise mutual information (PMI) by framing conditional probability estimation as a contrastive task with an explicit OTHER category. The authors benchmark PromptNCE against four other prompting‑based estimators on three human‑annotated datasets, finding that PromptNCE achieves the best conditional probability estimates and Spearman correlations up to 0.78 for full PMI. A case study demonstrates the method’s utility for scoring student knowledge summaries in low‑data settings, and the authors release code and prompts for reproducibility.

By Juliette Woodrow, Chris Piech