Hugging Face Trending Papers

TrustMI: Causally controlling how assistants trust their users

Read the original on Hugging Face Trending Papers →

The paper introduces TrustMI, a method to causally control how large language model assistants decide to trust their users. By creating 2,000 contrastive conversations that vary in ability, benevolence, and integrity, the authors learn steering matrices that adjust trust decisions along linear directions in model activations while keeping the model parameters frozen. Experiments across six instruction‑tuned models show that these steering changes reliably alter trust decisions and affect safety‑related behaviors such as compliance with harmful requests, prompt injections, and insider threats.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 17

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

The paper introduces "cunning questions"—non‑safety prompts that contain misleading premises or subtle inconsistencies—to train large language models (LLMs) to scrutinize underlying intent and assumptions. Experiments show that incorporating these questions improves robustness against out‑of‑distribution jailbreak attacks and enhances subsequent safety fine‑tuning, achieving a new state‑of‑the‑art reduction in mean ASR from 17.40% to 15.05% across nine backbone–benchmark combinations. The authors argue that this training fosters vigilance, enabling models to prioritize safety judgments before engaging in harmful planning.

By Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi
arXiv Machine Learning
Jun 2

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

arXiv:2606. 00566v1 Announce Type: new Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party content, their attack surface expands well beyond what users type.

By Mohammed Sameer Syed (University of Arizona), Rozhin Yasaei (University of Arizona)