The paper investigates how increasing inference-time computation—via wider beam search or sample‑plus‑vote—affects performance on grammar‑constrained text‑to‑SQL tasks for small language models. Using the Qwen2.5‑Instruct family (0.5B–7B parameters) on the Spider benchmark, the authors find that larger models consistently outperform higher inference compute on the same model size, and that beam search yields better accuracy than sample‑plus‑vote under matched budgets. These results suggest that, unlike unconstrained settings, scaling inference compute does not compensate for smaller model size when strict grammar constraints are applied.
By Ty Chermsirivatana, John MacCormick
arXiv:2608. 10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages.
By Wajdi Ben Saad, Safa Madiouni
arXiv:2609.38006v1 Announce Type: new
Abstract: Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable tha...
By Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran
The study investigates why language models exhibit systematic performance gaps across English dialects, a phenomenon termed the "dialect tax." Using parallel dialect corpora that preserve meaning while altering surface form, the authors confirm that models treat Standard American English and dialectal texts as semantically equivalent, yet find representational disparities that persist through tokenization, pre‑training, post‑training, and inference. Even a character‑level tokenizer does not eliminate input/output asymmetries or accuracy gaps, and dialect pairs produce more divergent gradient updates than unrelated Standard texts, indicating that dialectal content is harder for models to learn.
By Elle
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions.