The paper critiques a recent NLI benchmark that tests the imperfective paradox, arguing that the benchmark suffers from conceptual and evaluation mis-specifications, notably Aspectual Reduction and a lack of strict NLI standards. The authors re-evaluate the benchmark, identify mis-specifications, and construct lexically matched minimal pairs to control for lexical variation. Their experiments reveal that models often exhibit a Sufficiency Bias, accept simple‑past hypotheses without affirming culmination, and that prompting interventions shift label decisions without improving true semantic understanding, highlighting additional failure modes such as compositional aspectual classification errors and surface‑form attraction.
By Kaiqiao Han, Yizhou Sun
arXiv:2609.37121v1 Announce Type: new
Abstract: Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by en...
By Nikitas Theodoropoulos, Maria Lymperaiou, Giorgos Filandrianos
The paper reviews how large language models (LLMs) are employed to study infant syntax acquisition, focusing on the BabyLM challenge that aims for human‑level syntactic performance using developmentally realistic corpora. It critically examines dataset construction, model selection, training procedures, and syntactic evaluation methods, highlighting methodological assumptions that limit the theoretical reach of these studies. The authors find that using developmentally realistic corpora has only modest impact on benchmark performance, pointing to fundamental computational differences between LLMs and actual infant syntax learners.
By H\'elie Bazin (SCAI, SND, ISIR), Anouk Barberousse (SND), Fran\c{c}ois Yvon (MLIA)
arXiv:2509. 15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences.
By Th\'eo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin
arXiv:2606. 01134v1 Announce Type: cross Abstract: Automatically distinguishing child-directed speech from adult-directed speech in long-form recordings is key to scalable analyses of children's language environments.
By Th\'eo Charlot, Tarek Kunze, Kaveri K. Sheth, Alejandrina Cristia, Marvin Lavechin
The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.
By Yixuan Wang, Freda Shi, Kanishka Misra
arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.
By Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
arXiv:2606. 00230v1 Announce Type: new Abstract: Grokking, the phenomenon in which neural networks generalize long after fitting their training data, has been studied in supervised settings on many epochs.
By Sherin Muckatira, Namrata Shivagunde, Vijeta Deshpande, Anna Rumshisky
Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language.
arXiv:2608.22452v1 Announce Type: new
Abstract: Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it...
By Francisco Portillo L\'opez
arXiv:2609.01151v1 Announce Type: new
Abstract: In the standard LM training pipeline, subword tokenisation is applied as a preprocessing step. Subword segmental language modelling is an alternative p...
By Francois Meyer
arXiv:2609. 17435v1 Announce Type: new Abstract: We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.
By Adam Zachary Wasserman, David Beauchemin