arXiv Machine Learning By Yihe Deng, Yu Yang, Junkai Zhang, Wei Wang, Bo Li

Enhancing LLM Safety Through a Theoretical Minimax Game Lens

Read the original on arXiv Machine Learning →

arXiv:2502. 05163v2 Announce Type: replace-cross Abstract: The rapid advancement of large language models (LLMs) necessitates effective mechanisms to ensure their responsible deployment by accurately distinguishing unsafe content from benign content.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 16

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.

By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
arXiv AI
Aug 26

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

RePolicy is a reinforcement learning approach designed to invoke safety policies for language model agents by evaluating entire execution trajectories within context-dependent policy libraries. It generates policy-grounded rationales and safety judgments, and is initialized with the PolicyTraj-20K dataset before fine-tuning via GRPO with verifiable rewards and policy-context perturbation. Experiments on six safety benchmarks demonstrate strong safety-detection performance and robust policy invocation across varying contexts.

By Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He
arXiv AI
4d ago

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

The paper introduces OTROPE, a likelihood‑free method for off‑policy evaluation of large language models (LLMs) that uses optimal transport to align labeled samples from a behavior model with unlabeled samples from a target model in a semantic space. OTROPE corrects human‑labeled residuals with proxy predictors, achieving a doubly robust evaluation without requiring behavior‑policy modeling or density‑ratio estimation. The authors provide theoretical guarantees for consistency and convergence, and demonstrate through synthetic and real LLM tasks that OTROPE outperforms existing baselines and can elevate weaker evaluators to match or exceed stronger ones.

By Liner Xiang, Wenbo Zhang, Hengrui Cai