The paper investigates how increasing inference-time computation—via wider beam search or sample‑plus‑vote—affects performance on grammar‑constrained text‑to‑SQL tasks for small language models. Using the Qwen2.5‑Instruct family (0.5B–7B parameters) on the Spider benchmark, the authors find that larger models consistently outperform higher inference compute on the same model size, and that beam search yields better accuracy than sample‑plus‑vote under matched budgets. These results suggest that, unlike unconstrained settings, scaling inference compute does not compensate for smaller model size when strict grammar constraints are applied.
By Ty Chermsirivatana, John MacCormick
arXiv:2608. 10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages.
By Wajdi Ben Saad, Safa Madiouni
arXiv:2609.38006v1 Announce Type: new
Abstract: Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable tha...
By Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran
The study investigates why language models exhibit systematic performance gaps across English dialects, a phenomenon termed the "dialect tax." Using parallel dialect corpora that preserve meaning while altering surface form, the authors confirm that models treat Standard American English and dialectal texts as semantically equivalent, yet find representational disparities that persist through tokenization, pre‑training, post‑training, and inference. Even a character‑level tokenizer does not eliminate input/output asymmetries or accuracy gaps, and dialect pairs produce more divergent gradient updates than unrelated Standard texts, indicating that dialectal content is harder for models to learn.
By Elle
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions.
The paper investigates how multilingual large language models can be guided to reason more reliably in low- to mid-resource languages by selecting appropriate language modes during inference. Experiments with LLaMA and Qwen models show that using English context can correct errors from non‑English comprehension, but adding redundant bilingual context can cause interference. To balance this trade‑off, the authors propose Reliability‑Aware Adaptive Inference (RAAI), a training‑free test‑time framework that routes prompts based on Expected Calibration Error and gates reasoning with a mid‑layer Risk Index, achieving up to 37.7% accuracy gains and reduced calibration error on low‑resource languages.
By Ekata Mitra, Ameeta Agrawal
arXiv:2606. 26836v1 Announce Type: new Abstract: Existing benchmarks typically report accuracy for a single model on a single run.
By Bradley Fowler, Ryan Smith, Daniel Thi Graviet, William Myers, Joshua Greaves, Narmeen Fatimah Oozeer, Ant\'ia Garc\'ia, Philip Quirke, Amirali Abdullah, Fazl Barez, Shriyash Kaustubh Upadhyay
arXiv:2609.15079v1 Announce Type: cross
Abstract: Multi-agent LLM architectures, such as LangChain and AutoGen, largely assume English as the lingua franca for internal inter-agent communication, eve...
By Kushagra Agrawal, Yuming Feng, Man-Fai Leung
arXiv:2607. 25292v1 Announce Type: new Abstract: Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution.
By Chaemin Jang, Dongman Lee, Jihee Kim
arXiv:2605. 02608v2 Announce Type: replace-cross Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood.
By Kevin Guan, Happy Buzaaba, Christiane Fellbaum
arXiv:2609.14144v1 Announce Type: cross
Abstract: A transformer language model is trained to respond to any prompt, but each deployment asks only a narrow range of questions: a support assistant sees...
By Jerry Kaplan