arXiv AI By Zonghuan Xu, Xiang Zheng, Yutao Wu, Xingjun Ma

What Does a Sharing Question Add? Auditing LLM Survey Scores for Misinformation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Sep 18

An Analysis of Training-Free Self-Reported Confidence in Language Models

The paper investigates whether language models’ self-reported confidence is meaningful without additional training. By evaluating three training‑free signals—direct verbalization, post‑hoc probability estimates, and agreement across multiple generations—on 100 TriviaQA questions, the authors find that direct verbalization alone achieves high AUROC scores (0.956 and 0.937) for correctness prediction, while agreement-based methods perform noticeably worse. Re‑eliciting confidence for the same answers shows modest score shifts and occasional decision flips, and an audit of biography claims reveals only a small confidence gap between supported and contradicted statements.

By Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li
arXiv AI
Sep 15

Synthetic Data in Marketing Research: How to Evaluate and When to Trust

The paper discusses the use of synthetic data in marketing research, arguing that the key question is not whether synthetic respondents work, but when they do. It categorizes synthetic data into three types—ungrounded LLM responses, segment-level personas, and individual-level digital twins—and maps each to the decisions they can support. The authors also propose a taxonomy of accuracy measures, highlight the forgotten question problem, and introduce an ex‑ante answerability diagnostic based on R² to improve twin-human correlation.

By Oded Netzer, Rajan Sambandam
arXiv AI
Sep 4

MIRA: A Bilingual Benchmark for Medical Information Response Audit

MIRA is a bilingual benchmark that evaluates whether large language models (LLMs) provide consistent medical information across different user phrasings, languages, and health literacy levels. It contains 4,320 prompts derived from 60 medically reviewed low‑risk health questions and reveals that models tend to omit key information and offer fewer concrete next steps when responding to low health‑literacy signals, a phenomenon termed Differential Information Dilution (DID). A knowledge‑guided mitigation prompt can reduce this dilution for most models, notably improving Claude and Qwen.

By Mengyu Xu, Qiaoxin Yang, Qianqian Wang, Xiwei Dai, Weiyi Wu, Chongyang Gao