The paper investigates whether reader-specific differences in retrieval‑augmented generation (RAG) reflect reusable structure or merely input‑local interactions. By fixing query, evidence, task, scoring, and intervention, the authors find that nine readers disagree on the effect sign in 33% of cases, with reader×query interactions explaining 29.8% of utility variance. They further decompose heterogeneity into evidence activity, ordinal preference, and conditional signed direction, discovering that ordinal reader geometry is stable across multiple settings while signed geometry is task‑bounded, yet stable ordinal similarity does not predict cross‑reader intervention transfer.
arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.
By Valentin No\"el
arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.
By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
By Koyar Afrasyab
The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.
By Saad Aamir, Muhammad Awais Bin Adil
arXiv:2607. 28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias.
By Plawan Kumar Rath
arXiv:2608. 15428v1 Announce Type: cross Abstract: Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question.
By Volodymyr Ovcharov
The article examines how a difference‑in‑differences (DiD) analysis on a censored rating scale can produce misleading effects. It demonstrates that each DiD component is censored by its own share, causing differential attenuation that can fabricate an interaction effect when the two responses are unequally censored. Using a pre‑registered audit of an LLM judge, the authors show that the reported significant interaction is largely an artifact of this censoring mechanism, with the true preference effect being null.
By Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang
arXiv:2606. 20205v1 Announce Type: new Abstract: Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles that affect their usability, safety assessment, and use as proxies for human participants in research.
By Jelena Meyer, David Garcia, Dirk U. Wulff
arXiv:2607. 28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety.
By Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu
arXiv:2608.31017v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 1...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2606. 09843v3 Announce Type: replace-cross Abstract: Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behave.
By Juan Manuel Contreras