arXiv Computation and Language

WildSEEK: Evaluating Language Models for Information-Seeking

WildSEEK is a new dataset of 3,000 real user information‑seeking queries, manually annotated for risk‑sensitive domains and whether the query is factoid or analytical. The accompanying evaluation framework tests LLM responses against four failure criteria—sycophantic behavior, overreliance, a default US‑centric perspective, and poor handling of vulnerable populations—finding higher failure rates for analytical queries. The authors also train classifiers on WildSEEK to analyze over 1.8 million realistic queries, revealing that more than a third are high‑risk and often analytical.

arXiv AI
Jun 6

Evaluating the Utility of Personal Health Records in Personalized Health AI

arXiv:2605. 18937v2 Announce Type: replace Abstract: Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex, potentially hindering insights.

By Rory Sayres, Kejia Chen, Ayush Jain, Matthew Thompson, Jonathan Richina, Xiang Yin, Jimmy Hu, Fan Zhang, Bob Lou, Mike Sanchez, Ines Mezerreg, Meredith Schreier, Hamsa Subramaniam, I-Ching Lee, Yugang Jia, Daniel Mcduff, Yossi Matias, Avinatan Hassidim, Dale Webster, Yun Liu, Jackie Barr, Quang Duong
arXiv AI
3d ago

Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

The paper evaluates how modern large language models use internal web search to answer factual questions. Using 783 static queries and 288 dynamic queries, the authors find that enabling retrieval improves accuracy on static questions but hurts confidence calibration. On dynamic queries, models often retrieve but still achieve less than 70% accuracy, mainly due to poor query formulation and source selection, indicating that internal web search works better as a quick verification tool than a full information‑retrieval system.

By Sahil Kale
arXiv AI
Aug 26

Ad Insertion in LLM-Generated Responses

arXiv:2601.19435v2 Announce Type: replace-cross Abstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on s...

By Shengwei Xu, Zhaohua Chen, Xiaotie Deng, Zhiyi Huang, Grant Schoenebeck
arXiv Machine Learning
Jul 17

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

arXiv:2508. 00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against.

By Jiazhen Pan (Cherise), Bailiang Jian (Cherise), Paul Hager (Cherise), Yundi Zhang (Cherise), Che Liu (Cherise), Friederike Jungmann (Cherise), Hongwei Bran Li (Cherise), Julian Canisius (Cherise), Chenyu You (Cherise), Junde Wu (Cherise), Jiayuan Zhu (Cherise), Fenglin Liu (Cherise), Yuyuan Liu (Cherise), Niklas Bubeck (Cherise), Moritz Knolle (Cherise), Chen (Cherise), Chen (Cherise), Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert
arXiv Computation and Language
Aug 27

The "Curse of Knowledge" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion

The paper introduces concept provenance, a framework that categorizes query concepts into backstory-supported, human-central, human-tail, and candidate answer-side zones to detect knowledge intrusion in LLM-generated search queries. Applying this framework to 77,004 queries across 100 topics, 8 LLMs, and 5 prompt conditions, the authors find that 7.40% of non-generic concepts are candidate answer-side and appear in 97 of 100 topics, with human validation revealing 68.2% relaxed precision and equal contributions from knowledge and deployment intrusion. The study shows that while prompt conditions do not eliminate intrusion, post-generation concept-provenance selection can achieve 99% elimination, positioning concept provenance as a diagnostic tool rather than an evaluation predictor.

By Chenglong Ma, Xinye Wanyan, Danula Hettiachchi, Ziqi Xu, Jeffrey Chan
arXiv Computation and Language
2d ago

Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries

arXiv:2609.00319v1 Announce Type: cross Abstract: Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compos...

By Phuong Anh Nguyen, Jill Noorily, Matthew Flathers, Haruka Notsu, Laura Ospina-Pinillos, Tommy Nguyen, Samantha Clark, Aoife Keane, Grace Thompson, John Torous