arXiv:2607. 17883v1 Announce Type: cross Abstract: Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true.
By Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu
The paper introduces TrustMI, a method to causally control how large language model assistants decide to trust their users. By creating 2,000 contrastive conversations that vary in ability, benevolence, and integrity, the authors learn steering matrices that adjust trust decisions along linear directions in model activations while keeping the model parameters frozen. Experiments across six instruction‑tuned models show that these steering changes reliably alter trust decisions and affect safety‑related behaviors such as compliance with harmful requests, prompt injections, and insider threats.
arXiv:2609.13579v1 Announce Type: new
Abstract: Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understandin...
By Fanqi Zeng, Sadid A. Hasan, Chaocheng He
arXiv:2609.26035v1 Announce Type: new
Abstract: Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and e...
By Sebastian Cochinescu
The paper introduces Trustworthy RAG, an evaluation agent designed to detect misinformation and knowledge poisoning in Retrieval-Augmented Generation systems. It combines natural language inference verification, a five-signal poison detector, and a weighted Trust Index to assess the reliability of retrieved content. Experiments on multiple LLMs show high accuracy and precision, with the agent effectively blocking unsafe advice in a secure-coding assistant scenario.
By Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson
The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.
By Nyamtulla Shaik, Fengjun Li, Bo Luo