arXiv:2609.37121v1 Announce Type: new
Abstract: Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by en...
By Nikitas Theodoropoulos, Maria Lymperaiou, Giorgos Filandrianos
arXiv:2608. 13545v1 Announce Type: cross Abstract: Modern language models are trained on heterogeneous web-scale text corpora.
By Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd\"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
The study introduces a Difference in Surprisal method that uses GPT‑2 token surprisal to automatically label telicity in English CHILDES corpora, validated against expert judgments. Logistic regression classifiers trained on 12 syntactic and lexical semantic features reveal that child speech achieves near‑perfect telicity classification using a single deterministic cue—the presence of a post‑verbal determiner—whereas adult speech relies more on verb class and other lexical semantic features, with the determiner cue neutralized. This developmental trajectory supports syntactic bootstrapping, showing that learners initially exploit high‑frequency structural cues before developing fully compositional, verb‑based event structures.
By Ellie Xia, Parisa Kordjamshidi, Alan Hezao Ke
The study investigates how different types of caregiver feedback influence grammar learning by training small GPT‑2‑style models on child‑directed language and fine‑tuning them with reinforcement learning. Four feedback categories—communicative, structural alignment, semantic contingency, and affective—were evaluated, with structural alignment showing the strongest improvement in grammaticality and communicative feedback yielding moderate gains. Semantic contingency and affective feedback did not enhance grammaticality, though they may aid other language learning aspects, indicating that various feedback forms contribute complementarily to language acquisition.
LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.
By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv:2606.05087v2 Announce Type: replace
Abstract: Frequent verbs such as 'have' and 'make' can function either as collocates in light-verb constructions or as full lexical predicates, as in 'make a...
By Francesca Franzon, Nicolas Ros\`as G\'omez, Leo Wanner