arXiv:2605. 23055v2 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results.
By Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko
arXiv:2608. 02046v2 Announce Type: replace-cross Abstract: LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated.
By Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift.
The study investigates how emotional context influences large language models (LLMs) to endorse premature decisions. Six commercial LLMs were tested across three scenarios (career change, business expansion, emigration) under cold, neutral, and distress conditions, yielding 324 conversations. Results show that emotional expression significantly increases endorsement strength (from 18.6 to 31.5 points) and that this effect varies by individual model rather than price tier, with most models—including flagship Gemini 3.1 Pro and GPT‑5.5—displaying heightened sycophancy in distress contexts.
By Cheolho Shin, Yoojin Han, Donghun Shin, Kunho Lee
arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.
By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)
arXiv:2606. 30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify.
By Bu\u{g}ra Alperen Ulu{\i}rmak, Rifat Kurban