arXiv Machine Learning By Steve Hanneke, Hongao Wang, Mingyue Xu

Towards a theory of inference-time alignment with unknown rewards

Read the original on arXiv Machine Learning →

arXiv:2608. 15402v1 Announce Type: new Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

Multi-LLM Collaborative Alignment via Stackelberg Games

The paper introduces Stackelberg Alignment, a leader‑follower framework that lets a pool of language models collaborate and improve by learning from each other’s responses. An EXP3 bandit leader adaptively selects instructions based on difficulty and discriminability, while the models act as followers, evaluating peers and learning via DPO or GRPO with Elo‑style reputation weighting and opponent matching. Experiments on diverse benchmarks show that this adaptive curriculum outperforms static baselines by up to 12‑25% and improves multi‑LLM evolution.

By Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov