The paper surveys 61 studies on mental‑health AI and identifies a misalignment in how trust is evaluated across disciplines. It proposes a three‑layer framework—human‑oriented, interaction‑oriented, and AI‑oriented trust—and maps stakeholder perspectives onto these layers. The authors argue that future research should focus on calibrating human trust to actual interaction and AI trustworthiness rather than merely maximizing perceived trust.
By Xin Sun, Yue Su, Yifan Mo, Qingyu Meng, Yuxuan Li, Min Chen, Mengyuan Zhang, Saku Sugawara, Charlotte Gerritsen, Sander L. Koole, Koen Hindriks, Jiahuan Pei
arXiv:2607. 09586v1 Announce Type: new Abstract: The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them.
By Hannah M. Liu, Rhea Saxena, Shiv Asthana
arXiv:2606. 14923v1 Announce Type: new Abstract: As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates.
By Yujiao Chen
arXiv:2606. 30658v1 Announce Type: cross Abstract: Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outputs transparent to users.
By Zhiling Yan, Zhe Fang, David J King, Ann Pongsakul, Eashan Adhikarla, Hui Ren, Sunyang Fu, Quanzheng Li, Lifang He, Xiang Li, Hongfang Liu, Yonghui Wu, Lichao Sun
The article discusses autonomous systems as the pinnacle of AI development, emphasizing the need to blend connectionist and symbolic AI within systems engineering. It introduces a generic agent architecture that organizes behavior around long‑term memory and outlines challenges in linking sensory data to structured memory, goal‑oriented decision making, planning, and agent coordination for collective intelligence. The authors also explore agent trustworthiness, noting it extends beyond behavior to include cognitive validity, and propose methods for its evaluation while highlighting the gap between current capabilities and the envisioned autonomous multi‑agent systems.
By Joseph Sifakis
arXiv:2608.18265v2 Announce Type: replace-cross
Abstract: We introduce a general, easy-to-implement AI-based method for modeling and analyzing the structure and complexity of human behavior. We assig...
By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
arXiv:2607. 15992v1 Announce Type: new Abstract: Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings.
By Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji, Waheedullah Pardess, Jen Weedon, Jasmijn Remmers, Eliza Krigman, Matthew Ball, Yalda Daryani, Kiran Iqbal, Francielle Vargas, Mar\'ia Llorente S\'anchez, Joe Humphreys, Fendi Tsim, Kelly Fitzpatrick, Jeff Dunn, Catherine Feldman
The paper investigates whether trustworthiness scores and truth judgments produced by LLM-as-Judge systems are truly independent. Experiments on correctness-controlled QA show that trust scores align more closely with truth verdicts than human behavior does, indicating a weaker separation between the two. Stress tests that alter only the source attribution of identical QA pairs reveal that changes in trust scores also affect truth verdicts and associated probabilities, suggesting that trust scores should not be treated as independent evidence for truth judgments.
By Xin Sun, Di Wu, Yuchen Guo, Jiahuan Pei, Isao Echizen, Abdallah El Ali, Saku Sugawara
The paper presents an AI-based method that uses a large language model to emulate human decision-making by assigning it a "type vector" describing traits such as Altruism and Risk Aversion. By varying these dimensions and values, the authors fit the model to over 119,000 decisions from 78,657 participants in 10 classic economic games, finding that three dimensions—Risk Aversion, Strategic Sophistication, and Trust—sufficiently capture human behavior. The resulting type clusters, fewer than a dozen, predict behavior in new games, suggesting a low-dimensional, portable representation of human behavior across diverse settings.
By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
arXiv:2505. 23397v3 Announce Type: replace Abstract: This article presents a structured framework for Human-AI collaboration in Security Operations Centers (SOCs), integrating AI autonomy, trust calibration, and Human-in-the-loop decision making.
By Ahmad Mohsin, Helge Janicke, Ahmed Ibrahim, Iqbal H. Sarker, Seyit Camtepe
TrustFormer is a task‑specific framework that evaluates trust across multiple dimensions in dynamic collaborative systems. It synchronizes heterogeneous trust data using task identifiers and timestamps, then applies cross‑temporal and cross‑dimensional attention to model both temporal dynamics and inter‑dimensional correlations. By combining these multi‑dimensional trust profiles, the system selects optimal collaborators and achieves a 40.8% improvement in trust evaluation accuracy over existing methods.
By Botao Zhu, Xianbin Wang
The paper presents an AI-based method that models human behavior by assigning a language model a vector of trait intensities—called a type vector—and asking it to predict actions in various settings. By adjusting traits such as Altruism, Risk Aversion, Fairness, and Trust, the authors fit the model to 119,147 decisions from 78,657 subjects across 35 countries and 10 economic games, finding that three dimensions (Risk Aversion, Strategic Sophistication, and Trust) closely match human choices. The resulting type vectors cluster into fewer than a dozen groups and can predict behavior in new games with different rules, demonstrating the method’s generalizability and interpretability.
By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei