arXiv Machine Learning By Haolong Hu, Hanyu Li, Tiancheng He, Huahui Yi, An Zhang, Qiankun Li, Kun Wang, Yang Liu, Zhigang Zeng

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

Read the original on arXiv Machine Learning →

SaFeR-Steer is a progressive multi‑turn alignment framework that uses staged synthetic bootstrapping and tutor‑in‑the‑loop GRPO to train a single student model against adaptive, on‑policy attacks. It introduces Trajectory‑Consistent Summative Reward (TCSR) to ensure that low‑quality turns impact the overall trajectory reward. The authors release the STEER dataset and demonstrate that SaFeR‑Steer significantly improves safety and helpfulness on both single‑turn and multi‑turn benchmarks for Qwen2.5‑VL models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 4

Towards Multi-modal Multi-turn Safety: From Agentic Interaction to Strategic Alignment

The paper introduces MINT‑Safe, a new open‑source dataset of 11,270 multi‑image dialogues and 500 refusal VQA pairs designed to expose safety risks in multi‑modal large language models during open‑ended conversations. It also proposes TAD‑Align, a turn‑aware dual‑objective reward framework that dynamically up‑weights dialogue turns with inconsistent safety behavior, improving safety metrics on Qwen2.5‑VL‑7B‑Instruct and LLaVA‑Next‑7B. The results show over 10% reduction in attack success rate and notable gains in harmlessness and helpfulness while maintaining overall model performance.

By Han Zhu, Jiale Chen, Chengkun Cai, Shengjie Sun, Haoran Li, Yujin Zhou, Chi-Min Chan, Pengcheng Wen, Lei Li, Yike Guo, Sirui Han