arXiv AI

A Mental Model Based Framework of Trust

The paper proposes a mental model-based framework of trust that captures multidimensional aspects of trust and can be used to infer human trust in AI agents. It formalizes trust perception, trust evolution, human reliance, and decision-making, and defines the appropriate level of trust in the agent. Human subject studies evaluate whether adjusting human beliefs about the agent, as predicted by the framework, can change trust perceptions across performance, process, and purpose dimensions.

arXiv Computation and Language
Aug 24

Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers

The paper surveys 61 studies on mental‑health AI and identifies a misalignment in how trust is evaluated across disciplines. It proposes a three‑layer framework—human‑oriented, interaction‑oriented, and AI‑oriented trust—and maps stakeholder perspectives onto these layers. The authors argue that future research should focus on calibrating human trust to actual interaction and AI trustworthiness rather than merely maximizing perceived trust.

By Xin Sun, Yue Su, Yifan Mo, Qingyu Meng, Yuxuan Li, Min Chen, Mengyuan Zhang, Saku Sugawara, Charlotte Gerritsen, Sander L. Koole, Koen Hindriks, Jiahuan Pei
arXiv AI
Jul 1

Agentic AI Enhances Physician Trust in Clinical Decision Making

arXiv:2606. 30658v1 Announce Type: cross Abstract: Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outputs transparent to users.

By Zhiling Yan, Zhe Fang, David J King, Ann Pongsakul, Eashan Adhikarla, Hui Ren, Sunyang Fu, Quanzheng Li, Lifang He, Xiang Li, Hongfang Liu, Yonghui Wu, Lichao Sun
arXiv AI
6d ago

Bringing AI to Autonomous Systems -- From Cognition to Collective Intelligence

The article discusses autonomous systems as the pinnacle of AI development, emphasizing the need to blend connectionist and symbolic AI within systems engineering. It introduces a generic agent architecture that organizes behavior around long‑term memory and outlines challenges in linking sensory data to structured memory, goal‑oriented decision making, planning, and agent coordination for collective intelligence. The authors also explore agent trustworthiness, noting it extends beyond behavior to include cognitive validity, and propose methods for its evaluation while highlighting the gap between current capabilities and the envisioned autonomous multi‑agent systems.

By Joseph Sifakis
arXiv AI
Jul 20

Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI

arXiv:2607. 15992v1 Announce Type: new Abstract: Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings.

By Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji, Waheedullah Pardess, Jen Weedon, Jasmijn Remmers, Eliza Krigman, Matthew Ball, Yalda Daryani, Kiran Iqbal, Francielle Vargas, Mar\'ia Llorente S\'anchez, Joe Humphreys, Fendi Tsim, Kelly Fitzpatrick, Jeff Dunn, Catherine Feldman
arXiv AI
Aug 24

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

The paper investigates whether trustworthiness scores and truth judgments produced by LLM-as-Judge systems are truly independent. Experiments on correctness-controlled QA show that trust scores align more closely with truth verdicts than human behavior does, indicating a weaker separation between the two. Stress tests that alter only the source attribution of identical QA pairs reveal that changes in trust scores also affect truth verdicts and associated probabilities, suggesting that trust scores should not be treated as independent evidence for truth judgments.

By Xin Sun, Di Wu, Yuchen Guo, Jiahuan Pei, Isao Echizen, Abdallah El Ali, Saku Sugawara
arXiv AI
Aug 20

How AI Prompts Can Teach Us About the Structure of Human Behavior

The paper presents an AI-based method that uses a large language model to emulate human decision-making by assigning it a "type vector" describing traits such as Altruism and Risk Aversion. By varying these dimensions and values, the authors fit the model to over 119,000 decisions from 78,657 participants in 10 classic economic games, finding that three dimensions—Risk Aversion, Strategic Sophistication, and Trust—sufficiently capture human behavior. The resulting type clusters, fewer than a dozen, predict behavior in new games, suggesting a low-dimensional, portable representation of human behavior across diverse settings.

By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
arXiv Machine Learning
Aug 27

TrustFormer: Cross-Temporal and Cross- Dimensional Transformer for Task-Specific Multi-Dimensional Trust Evaluation

TrustFormer is a task‑specific framework that evaluates trust across multiple dimensions in dynamic collaborative systems. It synchronizes heterogeneous trust data using task identifiers and timestamps, then applies cross‑temporal and cross‑dimensional attention to model both temporal dynamics and inter‑dimensional correlations. By combining these multi‑dimensional trust profiles, the system selects optimal collaborators and achieves a 40.8% improvement in trust evaluation accuracy over existing methods.

By Botao Zhu, Xianbin Wang
arXiv AI
Sep 18

Modeling Human Behavior with Type Vectors Using AI

The paper presents an AI-based method that models human behavior by assigning a language model a vector of trait intensities—called a type vector—and asking it to predict actions in various settings. By adjusting traits such as Altruism, Risk Aversion, Fairness, and Trust, the authors fit the model to 119,147 decisions from 78,657 subjects across 35 countries and 10 economic games, finding that three dimensions (Risk Aversion, Strategic Sophistication, and Trust) closely match human choices. The resulting type vectors cluster into fewer than a dozen groups and can predict behavior in new games with different rules, demonstrating the method’s generalizability and interpretability.

By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei