SafeMath: Inference-time Safety improves Math Accuracy
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces ToxicGSM, a dataset of 1.9k arithmetic problems that embed harmful or sensitive context while keeping the math reasoning intact. It audits current large language models (LLMs) on this dataset, revealing how math word problems can subtly propagate bias, unethical, or psychologically harmful content, especially in educational settings for children. The authors propose SafeMath, a safety alignment technique that reduces harmful outputs without sacrificing, and sometimes improving, mathematical reasoning performance.
arXiv:2609.36254v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning bef...
arXiv:2506. 07031v5 Announce Type: replace-cross Abstract: Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities.
arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability.
The paper introduces "cunning questions"—non‑safety prompts that contain misleading premises or subtle inconsistencies—to train large language models (LLMs) to scrutinize underlying intent and assumptions. Experiments show that incorporating these questions improves robustness against out‑of‑distribution jailbreak attacks and enhances subsequent safety fine‑tuning, achieving a new state‑of‑the‑art reduction in mean ASR from 17.40% to 15.05% across nine backbone–benchmark combinations. The authors argue that this training fosters vigilance, enabling models to prioritize safety judgments before engaging in harmful planning.