arXiv AI By Guoxi Zhang, Jiawei Chen, Tianzhuo Yang, Lang Qin, Juntao Dai, Yaodong Yang, Jingwei Yi

Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry

Read the original on arXiv AI →

arXiv:2603. 26846v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

The paper introduces "cunning questions"—non‑safety prompts that contain misleading premises or subtle inconsistencies—to train large language models (LLMs) to scrutinize underlying intent and assumptions. Experiments show that incorporating these questions improves robustness against out‑of‑distribution jailbreak attacks and enhances subsequent safety fine‑tuning, achieving a new state‑of‑the‑art reduction in mean ASR from 17.40% to 15.05% across nine backbone–benchmark combinations. The authors argue that this training fosters vigilance, enabling models to prioritize safety judgments before engaging in harmful planning.

By Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi
arXiv AI
Aug 24

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

The paper identifies a vulnerability in large language models where harmful intent can be hidden within benign narratives, a phenomenon termed Semantic Camouflage. By examining latent activation patterns across several small language model families, the authors discover an "Intent Horizon"—a layer depth where harmful intent representations collapse. They propose Latent Intent Verification (LIV), a lightweight probing defense that detects harmful intent in early layers and outperforms existing guardrails on the PKU-SafeRLHF dataset.

By Md. Hasib Ur Rahman
arXiv AI
3d ago

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

The paper introduces CoT-Interpretability Alignment (CIA), a metric that quantifies how well a large language model’s chain-of-thought (CoT) explanations match its internal reasoning processes. Evaluated on two-hop question answering, hint intervention, and integer multiplication across three LLMs, the study finds limited alignment (44.8–75.9%) and demonstrates that post‑training with a reward combining task accuracy and parametric faithfulness can substantially improve CoT faithfulness without sacrificing accuracy. The authors provide a framework for auditing CoT faithfulness and a pathway to making explicit reasoning more trustworthy, with code and data publicly available.

By Yihuai Hong, Shauli Ravfogel, Chen Zhao, Eunsol Choi