arXiv AI By Myra Cheng, Lujain Ibrahim, Grace Liu, Michelle S. Lam, Vishakh Padmakumar, Nick Madibekov, Diyi Yang, Dan Jurafsky

LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 4

User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

The study investigates how users perceive the helpfulness and privacy-preservation of large language model (LLM) responses in privacy-sensitive scenarios. Using 94 participants and 90 PrivacyLens scenarios, researchers found that users’ evaluations of identical LLM outputs varied widely, whereas five proxy LLM judges showed high agreement but low correlation with user judgments. The results suggest that proxy LLMs cannot reliably estimate users’ diverse perceptions of utility and privacy, highlighting the need for more user-centered evaluation methods.

By Xiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer, Koichi Onoue