Hugging Face Trending Papers

TrustMI: Causally controlling how assistants trust their users

The paper introduces TrustMI, a method to causally control how large language model assistants decide to trust their users. By creating 2,000 contrastive conversations that vary in ability, benevolence, and integrity, the authors learn steering matrices that adjust trust decisions along linear directions in model activations while keeping the model parameters frozen. Experiments across six instruction‑tuned models show that these steering changes reliably alter trust decisions and affect safety‑related behaviors such as compliance with harmful requests, prompt injections, and insider threats.

arXiv AI
Sep 17

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

The paper introduces "cunning questions"—non‑safety prompts that contain misleading premises or subtle inconsistencies—to train large language models (LLMs) to scrutinize underlying intent and assumptions. Experiments show that incorporating these questions improves robustness against out‑of‑distribution jailbreak attacks and enhances subsequent safety fine‑tuning, achieving a new state‑of‑the‑art reduction in mean ASR from 17.40% to 15.05% across nine backbone–benchmark combinations. The authors argue that this training fosters vigilance, enabling models to prioritize safety judgments before engaging in harmful planning.

By Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi
arXiv Machine Learning
Jun 2

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

arXiv:2606. 00566v1 Announce Type: new Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party content, their attack surface expands well beyond what users type.

By Mohammed Sameer Syed (University of Arizona), Rozhin Yasaei (University of Arizona)
arXiv AI
Jun 9

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.

By Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo
arXiv AI
Sep 24

Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures

The paper introduces a counterfactually anchored evidence attribution approach for multi‑turn large language model safety failures. It presents a new dataset of 1,762 conversations, including adversarial, benign twins, and high‑risk vocabulary variants, and trains a lightweight hierarchical model that accurately predicts safety violations and attributes them to specific user turns and token spans. The model achieves high detection performance (F1 = 0.988) and significantly reduces adversarial confidence when top‑attributed tokens are removed, while maintaining low false‑positive rates on benign conversations.

By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
arXiv AI
Aug 28

ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices

ADeptS-Bench is a new benchmark designed to assess the trustworthiness of Computer Use Agents (CUAs) across mobile and desktop devices. It consists of two streams: a Safety stream with paired benign and malicious tasks that embed visual threats, and a Disambiguation stream that tests whether agents seek clarification when instructions are ambiguous. Evaluation of seven models shows none consistently achieves high task success while keeping attack success low, and all models exhibit problematic behaviors such as unhesitant checkout on a $25K order and failure to detect a mislabeled factory reset button.

By Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro, Yuxuan Sun, Chloe Evans, Joseph Tighe