arXiv AI

Language model agents show in-group trust bias invisible to standard behavioural audits

arXiv:2605. 28114v2 Announce Type: replace Abstract: Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and the same models increasingly control physically embodied robots as well as software.

arXiv Machine Learning
Aug 4

Trust or Check? Understanding the (Evolutionary) Dynamics of User Trust in AI Systems

arXiv:2603. 24742v2 Announce Type: replace-cross Abstract: As the capabilities and adoption of Artificial Intelligence (AI) systems grow, trust in these AI systems is an increasingly urgent concern.

By Adeela Bashir, Zhao Song, Ndidi Bianca Ogbo, Nataliya Balabanova, Martin Smit, Chin-wing Leung, Paolo Bova, Manuel Chica Serrano, Dhanushka Dissanayake, Manh Hong Duong, Elias Fernandez Domingos, Nikita Huber-Kralj, Marcus Krellner, Andrew Powell, Stefan Sarkadi, Fernando P. Santos, Zia Ush Shamszaman, Chaimaa Tarzi, Paolo Turrini, Grace Ibukunoluwa Ufeoshi, Victor A. Vargas-Perez, Alessandro Di Stefano, Simon T. Powers, The Anh Han
arXiv Machine Learning
Jun 9

Diffuse AI Control on Fuzzy Tasks

arXiv:2606. 08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment.

By Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton