arXiv Machine Learning By Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel

Which LLM to pick? Online Active Model Selection for Large Language Models

Read the original on arXiv Machine Learning →

The paper introduces ONLINE LLM PICKER, a framework for active model selection of large language models in streaming settings. It selects the most informative prompts for annotation within a limited budget, enabling the identification of the best or near‑best model among many candidates. Experiments on 10 datasets and over 130 language models show up to 71.67% savings in annotation cost and a reduction in regret by up to 2.51× when using the chosen model for sequential generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
6d ago

Large Language Model Selection with Limited Annotations

arXiv:2605.24981v2 Announce Type: replace Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotati...

By Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel
arXiv AI
Sep 3

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

The paper introduces SCX Router, a lightweight GLiClass-based model selector that assigns suitability scores to inference-time language models without autoregressive generation. It uses a 0.6B-parameter Qwen3 decoder with a shallow bidirectional scorer, preserving a text-only key–value cache across sessions and predicting task attributes such as type, difficulty, and expected output length. The authors build a comprehensive task ontology with 23 families, 115 types, and 1,173 synthetic examples, generating 150,000 verifier-scored tasks to train the router, which outperforms baseline models on LiveBench subsets with a top‑1 score of 0.707 versus 0.696 for the strongest fixed model.

By Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov