Hugging Face Trending Papers

Which Forms of Caregiver Feedback Support Grammar Learning? A Reinforcement-Learning Study of Child-Like Language Models

The study investigates how different types of caregiver feedback influence grammar learning by training small GPT‑2‑style models on child‑directed language and fine‑tuning them with reinforcement learning. Four feedback categories—communicative, structural alignment, semantic contingency, and affective—were evaluated, with structural alignment showing the strongest improvement in grammaticality and communicative feedback yielding moderate gains. Semantic contingency and affective feedback did not enhance grammaticality, though they may aid other language learning aspects, indicating that various feedback forms contribute complementarily to language acquisition.

arXiv Computation and Language
Sep 23

A retrospective analysis on the use of LLMs to study infant syntax learning

The paper reviews how large language models (LLMs) are employed to study infant syntax acquisition, focusing on the BabyLM challenge that aims for human‑level syntactic performance using developmentally realistic corpora. It critically examines dataset construction, model selection, training procedures, and syntactic evaluation methods, highlighting methodological assumptions that limit the theoretical reach of these studies. The authors find that using developmentally realistic corpora has only modest impact on benchmark performance, pointing to fundamental computational differences between LLMs and actual infant syntax learners.

By H\'elie Bazin (SCAI, SND, ISIR), Anouk Barberousse (SND), Fran\c{c}ois Yvon (MLIA)
arXiv Computation and Language
2d ago

Cross-Linguistic Effects in Bilingual Phoneme BabyLMs

arXiv:2609.37121v1 Announce Type: new Abstract: Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by en...

By Nikitas Theodoropoulos, Maria Lymperaiou, Giorgos Filandrianos
arXiv AI
1d ago

GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning

GrammarRL introduces a label‑free reinforcement learning approach that adapts language models to grammar constraints without annotated data. It optimizes two self‑supervised rewards—direct and reverse—using a Reinforce Leave‑One‑Out objective over grammar‑constrained rollouts, and regularizes toward a frozen base model. Experiments on sign‑language gloss translation, hierarchical text classification, and named entity recognition with Llama models show consistent gains over constrained greedy decoding and competitive performance to beam search while keeping inference cost low.

By Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov\`{\i}
Hugging Face Trending Papers
Jul 22

Exposure is Optional: Learning Unlike Coordination in Language Models

Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language.

arXiv Computation and Language
Sep 3

Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization

The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.

By Yixuan Wang, Freda Shi, Kanishka Misra