The paper reviews how large language models (LLMs) are employed to study infant syntax acquisition, focusing on the BabyLM challenge that aims for human‑level syntactic performance using developmentally realistic corpora. It critically examines dataset construction, model selection, training procedures, and syntactic evaluation methods, highlighting methodological assumptions that limit the theoretical reach of these studies. The authors find that using developmentally realistic corpora has only modest impact on benchmark performance, pointing to fundamental computational differences between LLMs and actual infant syntax learners.
By H\'elie Bazin (SCAI, SND, ISIR), Anouk Barberousse (SND), Fran\c{c}ois Yvon (MLIA)
arXiv:2609.01597v1 Announce Type: cross
Abstract: Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal struct...
By Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu
arXiv:2609.37121v1 Announce Type: new
Abstract: Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by en...
By Nikitas Theodoropoulos, Maria Lymperaiou, Giorgos Filandrianos
arXiv:2608. 10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives.
By Alec Harris, Kasey Corra, Archie Chaudhury, Yixiong Hao
GrammarRL introduces a label‑free reinforcement learning approach that adapts language models to grammar constraints without annotated data. It optimizes two self‑supervised rewards—direct and reverse—using a Reinforce Leave‑One‑Out objective over grammar‑constrained rollouts, and regularizes toward a frozen base model. Experiments on sign‑language gloss translation, hierarchical text classification, and named entity recognition with Llama models show consistent gains over constrained greedy decoding and competitive performance to beam search while keeping inference cost low.
By Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov\`{\i}
arXiv:2607. 06175v1 Announce Type: cross Abstract: Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data.
By Alexander Rombach, Chantale Lauer, Nijat Mehdiyev
Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language.
arXiv:2609. 17435v1 Announce Type: new Abstract: We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.
By Adam Zachary Wasserman, David Beauchemin
The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.
By Yixuan Wang, Freda Shi, Kanishka Misra
arXiv:2606. 09525v1 Announce Type: cross Abstract: During instruction fine-tuning (IFT), large language models (LLMs) learn to follow instructions by using the provided context to answer a query.
By Nadya Yuki Wangsajaya, Haeun Yu, Isabelle Augenstein
arXiv:2506. 10341v2 Announce Type: replace Abstract: Interactively learning from observation and language feedback is an increasingly studied area driven by the emergence of large language model (LLM) agents.
By Wanqiao Xu, Allen Nie, Ruijie Zheng, Aditya Modi, Adith Swaminathan, Ching-An Cheng
arXiv:2609.21231v1 Announce Type: new
Abstract: Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and t...
By Ruotian Wu, Bill E. Johnson, Gene Saunders, Osama Hamzeh, Ankit Vadehra, Pascal Poupart