arXiv AI

Decoupled Alignment for Robust Plug-and-Play Adaptation

arXiv:2406. 01514v4 Announce Type: replace-cross Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback.

arXiv AI
Jul 14

Stable On-Policy Distillation through Adaptive Target Reformulation

arXiv:2601. 07155v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models; however, conventional supervised KD often suffers from a distribution mismatch between training and inference.

By Ijun Jang, Jewon Yeom, Juan Yeo, Hyunggyu Lim, Taesup Kim
arXiv Computer Vision
Aug 28

LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

LLaVAFlow is an information‑theoretic distillation framework designed to preserve cross‑modal alignment in Multimodal Large Language Models during visual instruction tuning. It compresses the mutual information between extracted relations and MLLM embeddings to refine alignment flow, and then maximizes mutual information between pretrained and fine‑tuned alignment flows to transfer compact alignment information. Experiments demonstrate that LLaVAFlow effectively maintains alignment flow, improving downstream performance and generalization.

By Muyao Yuan, Muyan Jiao, Jiangyong Ying, Weizhan Zhang, Yuanhong Zhang, Lan Ma, Yuan Gao, Haipeng Du
arXiv Computation and Language
Sep 1

ACTD: Anchor-Based Cross-Tokenizer Distillation with Residual Regularization

The paper introduces ACTD, an Anchor-Based Cross-Tokenizer Distillation method that aligns vocabularies and sequences to transfer reasoning capabilities from large language models to smaller students. It uses a novel anchor loss with residual regularization to reduce alignment noise and extends the approach to multiple teachers. Experiments on five reasoning benchmarks with three teachers show state‑of‑the‑art results, with the multi‑teacher variant outperforming existing baselines.

By Huiyi Zhang, Zijian Li, Xiaocheng Feng, Weitao Ma, Xiaoliang Yang, Yichong Huang, Bing Qin
Hugging Face Trending Papers
Jun 10

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language. Existing inference-time defenses that mix logits from a safe anchor model require both models to share a vocabulary, which rules them out for the cross-family specialists where safety is most degraded.