arXiv Computation and Language

SafeMath: Safe Solutions for Unsafe Math Word Problems

The paper introduces ToxicGSM, a dataset of 1.9k arithmetic problems that embed harmful or sensitive context while keeping the math reasoning intact. It audits current large language models (LLMs) on this dataset, revealing how math word problems can subtly propagate bias, unethical, or psychologically harmful content, especially in educational settings for children. The authors propose SafeMath, a safety alignment technique that reduces harmful outputs without sacrificing, and sometimes improving, mathematical reasoning performance.

arXiv Computation and Language
Sep 1

SafeMath: Inference-time Safety improves Math Accuracy

arXiv:2603.25201v2 Announce Type: replace Abstract: Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-viola...

By Sagnik Basu, Subhrajit Mitra, Aman Juneja, Somnath Banerjee, Rima Hazra, Animesh Mukherjee
arXiv Computation and Language
Sep 24

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

SafeTutors is a benchmark designed to evaluate both safety and pedagogical effectiveness of AI tutoring systems across mathematics, physics, and chemistry. It introduces a risk taxonomy of 11 harm dimensions and 48 sub‑risks based on learning‑science literature, focusing on issues such as answer over‑disclosure, misconception reinforcement, and loss of scaffolding. The study finds that all tested models exhibit broad harms, that larger scale does not mitigate these issues, and that multi‑turn interactions significantly increase pedagogical failures from 17.7% to 77.8%.

By Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva, Sayan Layek, Somnath Banerjee, Julia Stoyanovich, Mykola Pechenizkiy
arXiv Machine Learning
Jul 24

Concept Concentration for Faithful Representation Intervention

arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.

By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han
arXiv AI
Sep 17

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

The paper introduces "cunning questions"—non‑safety prompts that contain misleading premises or subtle inconsistencies—to train large language models (LLMs) to scrutinize underlying intent and assumptions. Experiments show that incorporating these questions improves robustness against out‑of‑distribution jailbreak attacks and enhances subsequent safety fine‑tuning, achieving a new state‑of‑the‑art reduction in mean ASR from 17.40% to 15.05% across nine backbone–benchmark combinations. The authors argue that this training fosters vigilance, enabling models to prioritize safety judgments before engaging in harmful planning.

By Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi
arXiv Computation and Language
Aug 25

Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks

arXiv:2506.07356v3 Announce Type: replace Abstract: While Finetuning-as-a-Service (FaaS) enables customization of Large Language Models (LLMs) using user data, this service is vulnerable to safety de...

By Seokil Ham, Yubin Choi, Yujin Yang, Seungju Cho, Younghun Kim, Changick Kim
Hugging Face Trending Papers
Jun 2

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate computation to code. Yet models are still often used in settings where they must reason directly from natural language, and trustworthy models should solve small-number arithmetic word problems without external tools.

arXiv AI
Jul 7

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

arXiv:2607. 02914v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge.

By Jiyang Guan, Yong Xie, Jun Chen, Jiexi Liu, Zipeng Ye, Defeng Li, Jiayu Shen, Jialing Tao, Hui Xue