arXiv AI

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

arXiv:2606. 11201v1 Announce Type: cross Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions.

arXiv Machine Learning
Sep 22

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

arXiv:2609.24983v1 Announce Type: cross Abstract: We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction a...

By Lei Yang, Mengyin Liu, Jia Wang, Hangyu Guo, Liang Zhao, Zheng Ge, Kang An, Binxing Jiao, Qi Han, Daxin Jiang, Siqi Shen, Xiangyu Zhang
arXiv AI
Jul 17

Decoupled Alignment for Robust Plug-and-Play Adaptation

arXiv:2406. 01514v4 Announce Type: replace-cross Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback.

By Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu
arXiv AI
Sep 3

Automated Researchers Can Mitigate Well-characterized Alignment Failures

The paper investigates whether automated alignment researchers (AARs) can post‑train language models to reduce well‑characterized alignment failures such as deception, sycophancy, and jailbreaks while preserving general capability. Across ten failures, the strongest AAR methods significantly lower targeted failures and generalize to held‑out benchmarks, larger models, and multi‑turn audits. In contrast, a human baseline of 28 experienced researchers, given eight hours to devise one‑shot methods, underperformed the best AAR approaches, and providing human ideas to AARs did not improve results.

By Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
Hugging Face Trending Papers
Jun 10

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language. Existing inference-time defenses that mix logits from a safe anchor model require both models to share a vocabulary, which rules them out for the cross-family specialists where safety is most degraded.

arXiv AI
3d ago

Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts

Ready2Blend is a method that blends natural-language instructions with learned alignment prompts to enable continual alignment of large language models without retraining the backbone. It uses AlignFormer to map each requirement to a fixed-length prompt stored in a modular bank, while keeping the backbone and prior prompts frozen. The approach achieves 93.1–98.5% of joint‑training performance, retains prior knowledge, and reduces training time by up to 4.3×, also allowing weighted personalization and order‑free composition.

By Jeesu Jung, Hwan Chang, Juseon Do, Jeonghwan Choi, Jinho Choo, Sungwoo Nam, S. K. Hong, Hwanjun Song