arXiv Computation and Language
Sep 2

SafeMath: Safe Solutions for Unsafe Math Word Problems

The paper introduces ToxicGSM, a dataset of 1.9k arithmetic problems that embed harmful or sensitive context while keeping the math reasoning intact. It audits current large language models (LLMs) on this dataset, revealing how math word problems can subtly propagate bias, unethical, or psychologically harmful content, especially in educational settings for children. The authors propose SafeMath, a safety alignment technique that reduces harmful outputs without sacrificing, and sometimes improving, mathematical reasoning performance.

By Sagnik Basu, Subhrajit Mitra, Aman Juneja, Somnath Banerjee, Rima Hazra, Animesh Mukherjee
arXiv Machine Learning
Jul 24

Concept Concentration for Faithful Representation Intervention

arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.

By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han
arXiv AI
Sep 17

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

The paper introduces "cunning questions"—non‑safety prompts that contain misleading premises or subtle inconsistencies—to train large language models (LLMs) to scrutinize underlying intent and assumptions. Experiments show that incorporating these questions improves robustness against out‑of‑distribution jailbreak attacks and enhances subsequent safety fine‑tuning, achieving a new state‑of‑the‑art reduction in mean ASR from 17.40% to 15.05% across nine backbone–benchmark combinations. The authors argue that this training fosters vigilance, enabling models to prioritize safety judgments before engaging in harmful planning.

By Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi