arXiv Machine Learning By Jiahao Zeng, Ming Tang, Ningning Ding

Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 06178v1 Announce Type: new Abstract: Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 3

GMTRouter: Personalized LLM Router over Multi-turn User Interactions

GMTRouter is a personalized large language model router that represents multi‑turn user‑LLM interactions as a heterogeneous graph with five node types—user, LLM, query, response, and turn—to preserve relational structure. Using a lightweight inductive graph learning framework and a user‑conditioned graph sampling mechanism, it captures user preferences from few‑shot data, enabling effective personalization without extensive fine‑tuning. Experiments show GMTRouter outperforms strong baselines, improving accuracy by up to 0.108 and AUC by 0.124, and adapts to new users with minimal data.

By Yihang Sun, Encheng Xie, Tao Feng, Jiaxuan You
arXiv AI
Jun 4

Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling

arXiv:2606. 04284v1 Announce Type: cross Abstract: Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with human values.

By Yifan Wang, Jinyi Mu, Mayank Jobanputra, Yu Wang, Ji-Ung Lee, Soyoung Oh, Isabel Valera, Vera Demberg