Are Language Models Script-Aware?
arXiv:2610.08037v1 Announce Type: new Abstract: Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research ha...
The paper examines how large language models (LLMs) encode script knowledge across their layers using logistic regression probing and logit‑lens analysis. Findings show that both the input and instructed output scripts are represented in the earliest layers, but the model commits to the actual output script only in the final layers, with intermediate representations defaulting to Latin. This two‑stage process is consistent across methods and is more pronounced in larger models, suggesting a link between model depth and script commitment.
arXiv:2610.08037v1 Announce Type: new Abstract: Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research ha...
The paper "What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks" analyzes 14,767 arXiv submissions from 2022 to 2026 that introduce or update evaluation resources for large language models. It systematically maps changes in target systems, domains, evaluation materials, conditions, and scoring mechanisms, revealing a growing emphasis on action, interaction, and professional applications. The study also notes uneven development in model participation, with LLM-based scoring increasing in both agent and non-agent groups, while model-generated materials do not show a comparable rise.
arXiv:2606. 11166v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly described as performing at the level of human experts on knowledge economy tasks.
The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.
arXiv:2609.07876v1 Announce Type: cross Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...
arXiv:2507.21831v2 Announce Type: replace-cross Abstract: LLMs are seeing widespread use for task automation, including automated coding in the social sciences. However, even though researchers have...
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question.
arXiv:2606. 03618v1 Announce Type: new Abstract: AI-assisted coding agents are bottlenecked by input-token cost.
The paper proposes the interlingua hypothesis, suggesting that large language models translate by encoding a source sentence into a latent, task‑agnostic feature space and then decoding a target sentence from that space. Three lines of evidence support this: (1) BLEU variance across language pairs is largely explained by language‑specific competences without pair‑specific interactions; (2) many model components influence both monolingual and translation tasks; and (3) fine‑tuning on monolingual data recovers most translation gains seen with aligned documents. These findings converge to support the hypothesis and point toward new ways to understand and improve LLM translation.
arXiv:2607. 05188v1 Announce Type: new Abstract: A coding agent solving a software-engineering task spends dozens of steps reasoning, editing code, and running tests, yet little is known about what the underlying language model internally represents about the program it is working on.
arXiv:2607. 06008v3 Announce Type: replace Abstract: While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories.
arXiv:2508. 16131v3 Announce Type: replace-cross Abstract: Code completion entails the task of providing missing tokens given a surrounding context.