The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
By Titus von der Malsburg, Sebastian Pad\'o
arXiv:2609.01356v1 Announce Type: new
Abstract: Multilingual large language models (mLLMs) achieve strong performance in machine translation, yet our understanding of the mechanisms by which they tra...
By Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov, Josef van Genabith, Simon Ostermann
arXiv:2609.00416v1 Announce Type: new
Abstract: Probing studies have established that syntactic information is decodable in early and middle transformer layers, but what happens to that information i...
By Christos Nikolaos Zacharopoulos, Revekka Kyriakoglou, Chara Tsoukala, Th\'eo Desbordes
arXiv:2605. 02608v2 Announce Type: replace-cross Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood.
By Kevin Guan, Happy Buzaaba, Christiane Fellbaum
arXiv:2608.28924v1 Announce Type: new
Abstract: Linguistic theory has long recognized cross-linguistic syntactic regularities, leading to claims that these similar structures are processed by similar...
By Sasha Boguraev, Toshiki Nakai, Kyle Mahowald, Julius Steuer
arXiv:2606. 30815v1 Announce Type: cross Abstract: Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to be unacquirable by humans.
By Ram Janarthan, Coleman Haley, Sharon Goldwater
arXiv:2402. 18121v2 Announce Type: replace-cross Abstract: This study assesses four cutting-edge language models in the underexplored Aminoacian language.
By Yunze Xiao, Yiyang Pan
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
MURANO is an open‑source framework that enables researchers to design, run, and reproduce mechanistic interpretability experiments on large language models. It unifies the five key stages—loading, recording, attribution, intervention, and evaluation—into composable pipeline steps that exchange named artifacts and use canonical addresses for interoperability. The authors demonstrate the framework with reproductions of existing studies and a sparse autoencoder case study, showing its practical applicability across disciplines.
By Alireza Bayat Makou, Emirhan B\"oge, Phu Gia Hoang, Federico Tiblias, Jingcheng Niu, Subhabrata Dutta, Richard Eckart de Castilho, Iryna Gurevych
arXiv:2608.23448v1 Announce Type: new
Abstract: This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradi...
By Chit-Fung Lam
arXiv:2606. 05792v1 Announce Type: cross Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language still requires time and expertise, which limits adoption.
By Arslan Bisharat, Brian Ortiz, Eric Spencer, Khushboo Bhadauria, TaiNing Wang, George K. Thiruvathukal, Konstantin Laufer, Mohammed Abuhamad
The paper investigates how machine‑translated English data from 24 diverse source languages influences small English language models. It finds that source language affects model behavior: lexical diversity drives overall perplexity, while grammatical performance correlates with typological similarity to English when sufficient data is used. Additionally, translation quality strongly predicts language‑modeling performance.
By Jenny Kunz