arXiv AI

Whom to Query for What: Adaptive Group Elicitation via Multi-Turn LLM Interactions

arXiv:2602. 14279v2 Announce Type: replace-cross Abstract: Eliciting information to reduce uncertainty about latent group-level properties from surveys and other collective assessments requires allocating limited questioning effort under real costs and missing data.

arXiv AI
Sep 3

Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

The paper introduces ProSE, a framework for AI assistants that generate proposals while considering users’ bounded rationality and evaluability constraints. It proposes a KL‑regularised bounded‑rational binary response model and a depth‑2 Bayes‑adaptive planner, “ProSE‑Plan,” which scores proposals by expected responses and resulting belief updates. Experiments on graph simulations show that “ProSE‑Plan” outperforms evaluability‑unaware and myopic baselines, especially when evaluation cost is high, and that informative probes are crucial for effective assistance.

By Yifan Zhu, Sammie Katt, Samuel Kaski
arXiv AI
Aug 24

Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants

The paper introduces Evaluation-as-Search (EaS), a feedback‑driven method that adaptively probes LLM‑powered meeting assistants by focusing on natural questions likely to reveal grounding failures. Using EaS, the authors build MeetingProbe, a benchmark of over 3,000 annotated question‑answer pairs from 20 transcripts across three meeting genres and three assistants. Ablation studies show that adaptive search uncovers 2.5× more failures than random probing, revealing a capability gradient and eight recurring failure categories dominated by discourse‑pragmatic challenges.

By Sami Khairy, Yasaman Hosseinkashi, Vishak Gopal, Ross Cutler
arXiv Computation and Language
Sep 16

Towards Detecting AI-Assisted Responses in Online Surveys

The paper introduces ASURRE, a benchmark dataset for detecting AI‑assisted responses in online surveys. It evaluates how different LLM usage strategies—ranging from full generation to persona‑grounded agentic completion—affect the performance of existing machine‑generated text detectors. The study finds that while naive AI usage is easily detected, more sophisticated persona‑grounded agents approach chance performance, yet still leave identifiable behavioural traces that can be aggregated to improve detection.

By Qizhou Wang, Bogdan Mamaev, Christopher Leckie