arXiv Computation and Language By Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu, Amin Mantrach, Fabrizio Silvestri

QUORUM: QUality-Optimized Routing Using Multiple annotators

Read the original on arXiv Computation and Language →

QUORUM is a budget‑aware routing framework that dynamically assigns each annotation instance to either human or large language model (LLM) annotators, using feature‑based signals to estimate difficulty and multiple annotations per instance. It combines annotations through agreement‑based rewards to improve reliability, without relying on model confidence or uncertainty estimates. Experiments on diverse closed‑ and open‑ended tasks in English and multilingual settings show that QUORUM can boost annotation quality by up to 34.4% while cutting costs by 8.8% compared to existing methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 27

VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval

The paper introduces VDAR-Router, a routing framework for large language models that uses verbalized query difficulty analysis to guide model selection. It first generates an explicit difficulty profile for each query, retrieves historical examples with similar profiles, and then estimates model suitability to choose a model based on a reward function balancing performance and cost. Experiments on three datasets show that VDAR-Router consistently outperforms existing baselines in cost‑performance trade‑offs, and case studies confirm that explicit difficulty analysis improves example relevance and routing reliability.

By Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng, An-Zi Yen
arXiv AI
Aug 12

A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

arXiv:2608. 10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages.

By Wajdi Ben Saad, Safa Madiouni