arXiv Computation and Language

QUORUM: QUality-Optimized Routing Using Multiple annotators

QUORUM is a budget‑aware routing framework that dynamically assigns each annotation instance to either human or large language model (LLM) annotators, using feature‑based signals to estimate difficulty and multiple annotations per instance. It combines annotations through agreement‑based rewards to improve reliability, without relying on model confidence or uncertainty estimates. Experiments on diverse closed‑ and open‑ended tasks in English and multilingual settings show that QUORUM can boost annotation quality by up to 34.4% while cutting costs by 8.8% compared to existing methods.

arXiv Computation and Language
Aug 27

VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval

The paper introduces VDAR-Router, a routing framework for large language models that uses verbalized query difficulty analysis to guide model selection. It first generates an explicit difficulty profile for each query, retrieves historical examples with similar profiles, and then estimates model suitability to choose a model based on a reward function balancing performance and cost. Experiments on three datasets show that VDAR-Router consistently outperforms existing baselines in cost‑performance trade‑offs, and case studies confirm that explicit difficulty analysis improves example relevance and routing reliability.

By Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng, An-Zi Yen
arXiv AI
Aug 12

A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

arXiv:2608. 10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages.

By Wajdi Ben Saad, Safa Madiouni
arXiv AI
Sep 16

MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

MiCRo is a two‑stage framework that improves personalized preference learning for large language models. It first uses a context‑aware mixture model to capture diverse human preferences from large binary preference datasets, then applies an online routing strategy to dynamically adjust mixture weights based on context, reducing ambiguity. Experiments on multiple datasets show that MiCRo captures diverse preferences and enhances downstream personalization.

By Jingyan Shen, Jiarui Yao, Rui Yang, Yifan Sun, Feng Luo, Rui Pan, Tong Zhang, Han Zhao
arXiv AI
Jul 14

Agentic Routing: The Harness-Native Data Flywheel

arXiv:2607. 11399v1 Announce Type: cross Abstract: Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification.

By Xinchen Liu, Hang Zhou, Yingjie Zong, Yuchuan Tian, Liuyang Song, Shuo Zhang, Yulong Li, Wei He, Mengyu Zheng, Runke Liu, Siyang Cheng, Xiang Kuang, Hailin Hu, Kai Han, Yunhe Wang
arXiv AI
Sep 15

Towards Optimizing SQL Generation via LLM Routing

The paper "Towards Optimizing SQL Generation via LLM Routing" proposes a routing approach for Text-to-SQL tasks that dynamically selects the most cost‑effective large language model (LLM) for each query. Two routing strategies—score‑based and classification‑based—are introduced, achieving accuracy comparable to the best LLM while reducing latency and monetary cost. The authors design the routers for easy training and efficient inference, and demonstrate a practical accuracy‑cost trade‑off on the BIRD dataset.

By Mohammadhossein Malekpour, Nour Shaheen, Foutse Khomh, Amine Mhedhbi