Hugging Face Blog

Large Language Models: A New Moore's Law?

arXiv Computation and Language
Sep 17

Register Bias in Complexity-Based Large Language Model Routing

The paper examines how large language model (LLM) services route queries to models of varying size based on a cheap complexity estimate. It finds that this routing is not register neutral: queries written in non‑standard English registers (e.g., African American English or second‑language English) are systematically assigned to lower‑capacity models because they appear shorter due to omitted function words. Experiments on 37,704 learner sentence pairs and a controlled corpus show that this bias leads to significantly lower accuracy across all model tiers, including the highest‑capacity cloud models, while the routing decision itself adds little marginal cost.

By Simran Koul